Training, Validation, and Test Sets
The test set should not be used to guide model calibration.
Why Three Roles Matter
Model development is not only a train-then-test action. While developing a model, you make configuration choices and need feedback about those choices. Training data helps build the model, validation data helps guide development decisions, and test data provides a final evaluation after the model is ready.
The central rule is simple: use validation performance to guide configuration choices, but keep the test set out of that process. If test results influence model choices, the test set is no longer an independent final evaluation.
The Hold-Out Workflow
Simple hold-out validation assigns separate training, validation, and test roles. The training portion is used to build the model. The validation portion supplies feedback while you choose or adjust the model. The test portion remains aside until the model is ready for its final evaluation.
Choosing Between Two Configurations
A learner develops two model configurations and wants to decide which one to keep without compromising the final evaluation.
Build: Use the training set to build each configuration.
Compare: Use validation performance to compare the configurations and choose one.
Protect: Do not use the test results to decide between the configurations.
Evaluate: After the configuration is chosen, use the untouched test set once to estimate performance on new data.
Validation guides the choice; the test set supplies the independent final evaluation.
Rotating K-Fold Evaluation
K-fold validation addresses concern about split-dependent variation. The data is divided into K equal partitions. During each round, one partition takes the evaluation role while the remaining partitions are used for training. The evaluation role rotates until every partition has taken a turn. The resulting scores are then averaged.
Suppose the data is divided into four partitions. In the first round, partition 1 is evaluated and partitions 2, 3, and 4 are used for training. In the next round, partition 2 is evaluated and the other three partitions are used for training. The process continues until all four partitions have taken the evaluation role, after which the four scores are averaged.
Repeated Shuffling
When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each run, the data is shuffled and divided into K partitions again. Each run produces its own K-fold score, and the final score is the average of the scores from all runs.
Repeated K-fold validation requires training and evaluating P multiplied by K models when the process is repeated P times.Counting Repeated K-Fold Work
A validation procedure uses K folds and repeats the entire shuffled process P times.
Count folds in one run: One K-fold run trains and evaluates K models because each partition takes the evaluation role once.
Count repeated runs: Repeating the procedure P times creates P separate K-fold runs.
Combine the counts: The total number of trained and evaluated models is P multiplied by K.
Repeated shuffling gives evaluation exposure to more partition arrangements, but it increases computational cost.
Tuning Without Leaking the Test
Hyperparameters are configuration choices for a model. The source gives the number of layers and the size of layers as examples. They differ from model parameters such as network weights. When validation performance determines which configuration to keep, validation acts as a feedback signal in the search for a good configuration.
Repeated validation-based tuning can make the model fit the validation set indirectly. Each choice uses validation feedback, so the validation result is no longer a completely neutral observation of development. This is why the test set must remain untouched until the model is ready.
Mistakes in Evaluation Design
Using only training data and test data while choosing the model.
The test results have influenced model development, so the test set is no longer reserved for an independent final evaluation.
Fix:
Use a separate validation set to guide configuration choices and keep the test set untouched.Treating one hold-out split as unquestionably reliable on a small dataset.
Small validation and test portions may contain too few samples, and different shuffling and splitting rounds may produce very different measurements.
Fix:
Consider K-fold validation when split-dependent variation is a concern.Ignoring the cost of repeated K-fold validation.
Each repeated run contains K folds, so repeating the process requires training and evaluating P multiplied by K models.
Fix:
Use repeated K-fold validation when the added evaluation exposure justifies its computational cost.Treating validation performance as the final estimate of performance on new data.
Repeated validation-based tuning can make the model fit the validation set indirectly.
Fix:
After the model is ready, use the untouched test set once for final evaluation.
Apply the Workflow
A team has a small dataset. It wants to compare several configurations, is concerned that one random split may be unstable, and wants a final performance estimate that did not influence development. Describe a suitable evaluation workflow. State which data guides configuration choices, when K-fold validation could be used, and when the test set is evaluated.
Hints
- Separate development feedback from the final evaluation.
- K-fold validation rotates the evaluation role across partitions.
- If K-fold validation is repeated P times with K folds, count P multiplied by K model evaluations.
What do you think happens?
A developer checks the test score after every configuration change and keeps the configuration with the best test score. Is the final test score still an independent final evaluation?
Reveal answer
Answer: No, because test results guided configuration choices.
The test set must remain outside development. Once its results influence calibration or configuration choices, it no longer provides an untouched final evaluation.
Key Takeaways
- Training data builds the model, validation data guides development choices, and test data supplies the final evaluation.
- Hold-out validation keeps training, validation, and test roles separate.
- K-fold validation rotates the evaluation role across K partitions and averages the resulting scores.
- Repeated K-fold validation reshuffles and repartitions the data for multiple runs, requiring P multiplied by K model evaluations when repeated P times.
- Repeated tuning can fit the validation set indirectly, so the untouched test set must be used only after the model is ready.
Key Takeaways
- Use three dataset roles because building a model, guiding its development, and estimating final performance are different tasks.
- Use validation performance to choose configurations, not test performance.
- K-fold validation reduces dependence on one particular split by rotating evaluation across partitions.
- Repeated shuffled K-fold validation provides more partition arrangements but increases computational cost.
- Keep the test set untouched until the final evaluation because repeated validation-based tuning can overfit the validation set.