Model Calibration
The test set should not be used to guide model calibration.
The Protected Final Check
Model calibration means making choices or adjustments while developing a model. The central rule is that the test set should not guide those choices. If test results influence decisions about the model, the test set is no longer being held aside for an independent final evaluation.
Hold-Out Roles
The simplest validation protocol divides the available examples into three roles. Training data is used to build the model. Validation data is used while choosing or adjusting the model. Test data is held aside for the final performance measurement. This three-way separation is needed because a two-way division into training and test data is not enough when decisions are still being made.
A Three-Way Development Plan
Suppose a learner has a dataset and is still deciding between different model choices. Which role should each portion play?
Build: Use the training portion to build the candidate model.
Choose: Use the validation portion to compare or adjust candidate choices during development.
Protect: Do not use the test portion to guide those choices. Keep it aside for the final evaluation.
The validation portion supports calibration, while the test portion remains an independent final check.
Rotating K-Fold Evaluation
K-fold validation addresses some of the split-dependent variation of a single hold-out split. The data is divided into K equal partitions. One partition takes the validation role while the remaining partitions are used for training. The validation role then rotates so that each partition takes a turn. The resulting scores are averaged.
Four Partitions, Four Validation Turns
Imagine a K-fold process with four equal partitions.
Run 1: Partition 1 is used for validation, and the other three partitions are used for training.
Run 2: Partition 2 becomes the validation partition, while the other three are used for training.
Run 3: Partition 3 takes the validation role, and the remaining partitions are used for training.
Run 4: Partition 4 takes the validation role, and the remaining partitions are used for training.
Combine: Average the four resulting validation scores.
Every partition contributes a validation result, rather than relying on only one fixed validation portion.
Repeated Partitioning
When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each new run, the data is shuffled and divided into K partitions again. Each run therefore produces its own K-fold score. The final score is the average of the scores from all runs.
A single K-fold run explores one partition arrangement. Repeating the process exposes the evaluation to additional arrangements. This can help when performance measurements change substantially after different random shuffling and splitting choices.
The Cost of More Evidence
| Protocol | Partition behavior | Evaluation result | Computational implication |
|---|---|---|---|
| Simple hold-out | One fixed separation into training, validation, and test roles | One validation measurement and a final test measurement | Simplest protocol |
| K-fold validation | The validation role rotates across K equal partitions | Average of the K resulting scores | Requires training and evaluating K models for one K-fold run |
| Repeated K-fold validation | Shuffle and repartition before each K-fold run | Average the scores from all runs | Repeating the process P times requires training and evaluating P multiplied by K models |
The benefit of repeated validation has a direct cost. One K-fold run requires training and evaluating models for each of its K partitions. If the process is repeated P times, the total number of trained and evaluated models is P multiplied by K. Repetition can provide more exposure to different splits, but it requires more computation.
Mistakes in Calibration
Using the test score to choose or adjust the model
The test set has influenced the model decision, so it is no longer an independent final evaluation.
Fix:
Use a separate validation set to guide calibration and reserve the test set for the final measurement.Assuming a single hold-out split is always reliable
Small validation and test portions may not represent the broader data well.
Fix:
Treat instability across splitting rounds as a warning sign and consider K-fold validation.Confusing K-fold validation with one fixed validation split
That does not rotate the evaluation role across K partitions.
Fix:
Let each partition take a turn as validation, then average the resulting scores.Ignoring the cost of repeated K-fold validation
Each repetition performs another K-fold process.
Fix:
Remember that P repetitions with K partitions require training and evaluating P multiplied by K models.
Choose the Validation Protocol
A dataset is very small, and performance measurements change noticeably after different random shuffling and splitting rounds. The test set has not yet been used. Which validation approach is the strongest candidate, and what trade-off should be explained?
Hints
- Look for the method designed to respond to split-dependent variation.
- Consider what happens when the K-fold process is repeated several times.
Selecting a More Stable Evaluation Process
A small dataset produces unstable measurements under different hold-out splits. What process should be considered?
Identify the warning sign: Large changes between random splitting rounds indicate split-dependent variation.
Use rotating evaluation: K-fold validation allows more portions of the data to take part in evaluation by rotating the validation role.
Consider repetition: If the dataset is especially small and more precise evaluation is needed, shuffle and repartition before additional K-fold runs.
State the trade-off: Additional runs expose evaluation to more partition arrangements but require training and evaluating more models.
K-fold validation, possibly repeated with shuffling, addresses split-dependent variation at increased computational cost.
Key Takeaways
- Do not use test results to guide model calibration; reserve the test set for an independent final evaluation.
- Simple hold-out validation separates training, validation, and test roles.
- K-fold validation rotates the validation role across equal partitions and averages the resulting scores.
- Repeated K-fold validation reshuffles and repartitions the data before each run, then averages the run scores.
- Repeated validation can reduce dependence on one partition arrangement, but it increases computation because P repetitions with K partitions require P multiplied by K model runs.
Key Takeaways
- The test set must remain separate from calibration decisions.
- Hold-out validation gives training, validation, and test data distinct roles.
- K-fold validation rotates evaluation across K partitions and averages the scores.
- Repeated K-fold validation uses new shuffles and partitions for additional runs.
- More validation runs provide more split arrangements but increase computational cost.