K-Fold Cross-Validation and Model Selection
Try it: K-Fold Cross-Validation and Model Selection
How hold-out and k-fold cross-validation choose among candidate models: each fold takes a turn as the validation set, the fold errors are averaged, the lowest average picks the model, and only then is the untouched test set used once — while tuning on the test set would report an optimistic number.
How it works
- Set 10 test points aside; they take no part in any choice.
- Optionally shuffle the 20 development points (seeded), then cut them into k folds (hold-out: one chunk validates, the rest trains).
- For every round, train each candidate polynomial on the other folds and measure its validation MSE on the held-out fold.
- Average each candidate's fold errors (with repeated shuffling: average the p run scores; p × k fits per candidate).
- Select the candidate with the lowest average, retrain it on all development points, and evaluate it once on the test set.
- For contrast, show which candidate the test set itself would have picked — that figure is leaked and optimistic.
Default run (12 steps): 20 development points (sorted by x as they arrived) and 10 test points set aside. Candidates: polynomial degree 0, 1, 2, 3, 4, 5, 6. Plan: 5-fold cross-validation, shuffled with seed 1 → 5 fits per candidate (p × k = 1 × 5), 35 in total. … 5-fold cross-validation picks degree 3 (score 0.0768); its honest test MSE is 0.0889. Picking by test error would report degree 4 at 0.0868 — a leaked, optimistic figure.
Simplified: One fixed toy dataset (20 development + 10 test points from sin(2πx) with noise) and polynomial least squares of degree 0–6 as the candidate models. Shuffling uses a seeded mulberry32 Fisher–Yates; folds are contiguous chunks like scikit-learn's KFold. Hold-out uses the first of k chunks as the validation set.
Loading the simulation…