Concepts / Training, Validation, and Test Sets

Training, Validation, and Test Sets

The test set should not be used to guide model calibration.

  • Programming

Why Three Roles Matter

Model development is not only a train-then-test action. While developing a model, you make configuration choices and need feedback about those choices. Training data helps build the model, validation data helps guide development decisions, and test data provides a final evaluation after the model is ready.

The central rule is simple: use validation performance to guide configuration choices, but keep the test set out of that process. If test results influence model choices, the test set is no longer an independent final evaluation.

Training setBuild the modelValidation setGuide configurationTest setFinal evaluation
What is each dataset allowed to influence during model development?

The Hold-Out Workflow

Simple hold-out validation assigns separate training, validation, and test roles. The training portion is used to build the model. The validation portion supplies feedback while you choose or adjust the model. The test portion remains aside until the model is ready for its final evaluation.

assignedassignedheld asideAvailable dataExamplesTraining setBuild modelValidation setChoose configurationTest setUse once at the end
How are examples divided, and which set is used at each stage?

Choosing Between Two Configurations

A learner develops two model configurations and wants to decide which one to keep without compromising the final evaluation.

Build: Use the training set to build each configuration.

Compare: Use validation performance to compare the configurations and choose one.

Protect: Do not use the test results to decide between the configurations.

Evaluate: After the configuration is chosen, use the untouched test set once to estimate performance on new data.

Validation guides the choice; the test set supplies the independent final evaluation.

Rotating K-Fold Evaluation

K-fold validation addresses concern about split-dependent variation. The data is divided into K equal partitions. During each round, one partition takes the evaluation role while the remaining partitions are used for training. The evaluation role rotates until every partition has taken a turn. The resulting scores are then averaged.

rotaterotatecontinueaverage scoresRound 1P1 validationAverage scoreK resultsRound 2P2 validationRound 3P3 validationRound KPK validation
How does each partition take a turn as the validation set while the remaining partitions are used for training?

Suppose the data is divided into four partitions. In the first round, partition 1 is evaluated and partitions 2, 3, and 4 are used for training. In the next round, partition 2 is evaluated and the other three partitions are used for training. The process continues until all four partitions have taken the evaluation role, after which the four scores are averaged.

Repeated Shuffling

When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each run, the data is shuffled and divided into K partitions again. Each run produces its own K-fold score, and the final score is the average of the scores from all runs.

K-fold scoreK-fold scoreK-fold scoreRun 1Shuffle, K foldsFinal scoreAverage of P scoresRun 2Shuffle, K foldsRun PShuffle, K folds
What changes when the data is reshuffled and repartitioned for multiple K-fold runs?
Repeated K-fold validation requires training and evaluating P multiplied by K models when the process is repeated P times.

Counting Repeated K-Fold Work

A validation procedure uses K folds and repeats the entire shuffled process P times.

Count folds in one run: One K-fold run trains and evaluates K models because each partition takes the evaluation role once.

Count repeated runs: Repeating the procedure P times creates P separate K-fold runs.

Combine the counts: The total number of trained and evaluated models is P multiplied by K.

Repeated shuffling gives evaluation exposure to more partition arrangements, but it increases computational cost.

Tuning Without Leaking the Test

Hyperparameters are configuration choices for a model. The source gives the number of layers and the size of layers as examples. They differ from model parameters such as network weights. When validation performance determines which configuration to keep, validation acts as a feedback signal in the search for a good configuration.

buildselectevaluatefinal useTraining setBuild modelConfigurationKeep selected modelValidation setGuide choicesFinal evaluationNew-data estimateTest setUntouched
How does information flow during model selection while the test set stays isolated until the final evaluation?
comparechoose best repeatedlyCandidate modelsInitial choicesValidation setRepeated feedbackSelectedconfigurationRepeated validation choices
How can repeatedly choosing the best validation score make development fit the validation set indirectly?

Repeated validation-based tuning can make the model fit the validation set indirectly. Each choice uses validation feedback, so the validation result is no longer a completely neutral observation of development. This is why the test set must remain untouched until the model is ready.

Mistakes in Evaluation Design

  • Using only training data and test data while choosing the model.

    The test results have influenced model development, so the test set is no longer reserved for an independent final evaluation.

    Fix: Use a separate validation set to guide configuration choices and keep the test set untouched.

  • Treating one hold-out split as unquestionably reliable on a small dataset.

    Small validation and test portions may contain too few samples, and different shuffling and splitting rounds may produce very different measurements.

    Fix: Consider K-fold validation when split-dependent variation is a concern.

  • Ignoring the cost of repeated K-fold validation.

    Each repeated run contains K folds, so repeating the process requires training and evaluating P multiplied by K models.

    Fix: Use repeated K-fold validation when the added evaluation exposure justifies its computational cost.

  • Treating validation performance as the final estimate of performance on new data.

    Repeated validation-based tuning can make the model fit the validation set indirectly.

    Fix: After the model is ready, use the untouched test set once for final evaluation.

Apply the Workflow

MEDIUM

A team has a small dataset. It wants to compare several configurations, is concerned that one random split may be unstable, and wants a final performance estimate that did not influence development. Describe a suitable evaluation workflow. State which data guides configuration choices, when K-fold validation could be used, and when the test set is evaluated.

Hints
  • Separate development feedback from the final evaluation.
  • K-fold validation rotates the evaluation role across partitions.
  • If K-fold validation is repeated P times with K folds, count P multiplied by K model evaluations.

What do you think happens?

A developer checks the test score after every configuration change and keeps the configuration with the best test score. Is the final test score still an independent final evaluation?

  • Yes, because the test set was not used to build the model weights
  • No, because test results guided configuration choices
  • Yes, because the best score is the most informative score
Reveal answer

Answer: No, because test results guided configuration choices.

The test set must remain outside development. Once its results influence calibration or configuration choices, it no longer provides an untouched final evaluation.

Key Takeaways

  1. Training data builds the model, validation data guides development choices, and test data supplies the final evaluation.
  2. Hold-out validation keeps training, validation, and test roles separate.
  3. K-fold validation rotates the evaluation role across K partitions and averages the resulting scores.
  4. Repeated K-fold validation reshuffles and repartitions the data for multiple runs, requiring P multiplied by K model evaluations when repeated P times.
  5. Repeated tuning can fit the validation set indirectly, so the untouched test set must be used only after the model is ready.

Key Takeaways

  • Use three dataset roles because building a model, guiding its development, and estimating final performance are different tasks.
  • Use validation performance to choose configurations, not test performance.
  • K-fold validation reduces dependence on one particular split by rotating evaluation across partitions.
  • Repeated shuffled K-fold validation provides more partition arrangements but increases computational cost.
  • Keep the test set untouched until the final evaluation because repeated validation-based tuning can overfit the validation set.