Playground / K-Fold Cross-Validation and Model Selection

Pick a model without touching the test set

K-Fold Cross-Validation and Model Selection

Interactive lab

Try it: K-Fold Cross-Validation and Model Selection

How hold-out and k-fold cross-validation choose among candidate models: each fold takes a turn as the validation set, the fold errors are averaged, the lowest average picks the model, and only then is the untouched test set used once — while tuning on the test set would report an optimistic number.

How it works

  1. Set 10 test points aside; they take no part in any choice.
  2. Optionally shuffle the 20 development points (seeded), then cut them into k folds (hold-out: one chunk validates, the rest trains).
  3. For every round, train each candidate polynomial on the other folds and measure its validation MSE on the held-out fold.
  4. Average each candidate's fold errors (with repeated shuffling: average the p run scores; p × k fits per candidate).
  5. Select the candidate with the lowest average, retrain it on all development points, and evaluate it once on the test set.
  6. For contrast, show which candidate the test set itself would have picked — that figure is leaked and optimistic.

Default run (12 steps): 20 development points (sorted by x as they arrived) and 10 test points set aside. Candidates: polynomial degree 0, 1, 2, 3, 4, 5, 6. Plan: 5-fold cross-validation, shuffled with seed 1 → 5 fits per candidate (p × k = 1 × 5), 35 in total. … 5-fold cross-validation picks degree 3 (score 0.0768); its honest test MSE is 0.0889. Picking by test error would report degree 4 at 0.0868 — a leaked, optimistic figure.

Simplified: One fixed toy dataset (20 development + 10 test points from sin(2πx) with noise) and polynomial least squares of degree 0–6 as the candidate models. Shuffling uses a seeded mulberry32 Fisher–Yates; folds are contiguous chunks like scikit-learn's KFold. Hold-out uses the first of k chunks as the validation set.

Educational simulation

Loading the simulation…