Information Leakage
The test set should not be used to guide model calibration.
The Test Set Must Stay Independent
Information leakage occurs when information from the test set influences model calibration. Calibration includes choosing or adjusting the model. A test set is meant to provide an independent final evaluation, so using its results to guide decisions changes the role of that data.
Choosing Between Two Models
A learner builds two candidate models and uses performance measurements to decide which one to keep. Which dataset should guide that decision?
Separate the roles: Keep training data for building the model, use a separate validation set to guide the choice or adjustment, and reserve the test set for the final evaluation.
Make the calibration decision: Compare the candidate models using validation results rather than test results. The validation results are allowed to influence the model decision.
Evaluate once at the end: After the model has been chosen or adjusted, use the held-aside test set for the independent final evaluation.
The validation set guides calibration; the test set is reserved for the final evaluation.
Three Roles in Hold-Out Validation
Simple hold-out validation assigns separate roles to three portions of the data. The training portion is used to build the model. The validation portion is used while choosing or adjusting the model. The test portion is held aside for the final evaluation. This separation prevents calibration decisions from being guided by the data intended for the final check.
Hold-out validation is the simplest protocol, but it can be unreliable when only a small amount of data is available. The validation and test portions may contain too few samples to represent the broader data well. Instability is a warning sign: different random shuffling and splitting rounds produce very different performance measurements.
K-Fold Rotation
K-fold validation addresses split-dependent variation by dividing the data into K equal partitions. One partition takes the evaluation role while the remaining partitions are used for training. The evaluation role then rotates, so each partition serves as the evaluation portion once. The resulting scores are averaged.
A Four-Partition Rotation
A dataset is divided into four equal partitions. How does one K-fold validation run use those partitions?
Round 1: Partition 1 serves as the evaluation portion, while the other three partitions are used for training.
Round 2: Partition 2 serves as the evaluation portion, while the other three partitions are used for training.
Continue the rotation: Partitions 3 and 4 each take the evaluation role in turn, with the remaining partitions used for training in each round.
Combine the results: The scores from the four rounds are averaged to produce the K-fold score for this run.
Every partition contributes once as the evaluation portion, and the run produces an average score.
Use K-fold validation when a single hold-out split may be too dependent on which samples happened to land in the validation portion. It allows more portions of the data to take part in evaluation, while still rotating the training and evaluation roles.
Repeated Shuffling and Cost
When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each run, the data is shuffled and divided into K partitions again. Each run therefore uses a new partition arrangement, produces its own K-fold score, and contributes that score to the final average.
Repeated shuffling increases exposure to different partition arrangements, but it also increases computational cost. If the process is repeated P times and each run uses K folds, training and evaluating more models is required: P multiplied by K models in total.
Validation Versus Final Testing
| Aspect | Validation performance | Test performance |
|---|---|---|
| Role | Guide model choice or adjustment | Provide the final evaluation |
| May influence calibration? | Yes | No |
| Data status | Used during the decision process | Held aside for an independent final evaluation |
| If used incorrectly | The calibration process may be affected | The final estimate is no longer independent |
The difference is about direction of influence. Training data influences model building. Validation data can influence calibration. Test data should receive the final model without influencing the decisions that produced it. Confusing these directions is the central information-leakage mistake.
Mistakes That Cause Leakage
Using the test score to choose between candidate models
The test results have influenced model calibration, so the test set is no longer reserved for an independent final evaluation.
Fix:
Use a separate validation set to guide the choice, then use the test set for the final evaluation.Treating a training-and-test split as sufficient while still adjusting the model
The test results are guiding repeated calibration decisions.
Fix:
Introduce a validation role between model building and final testing.Assuming one hold-out split is always reliable
Small validation and test portions may not represent the broader data well, and instability signals split-dependent variation.
Fix:
Consider K-fold validation, or repeated shuffled K-fold validation when the dataset is especially small and evaluation precision is important.Repeating K-fold validation without accounting for computational cost
Each repeated run contains K training and evaluation rounds.
Fix:
For P repeated runs and K folds per run, account for training and evaluating P multiplied by K models.
Practice the Data Roles
A learner first builds a model with training data, then changes the model twice after looking at test performance. The learner reports the last test performance as the final result. Identify the leakage, name the missing data role, and describe how the process should be reorganized.
Hints
- Ask which dataset influenced the model changes.
- The final evaluation requires data that was not used to guide calibration.
- Name the separate role that should guide choosing or adjusting the model.
Practice Solution
Reorganize the learner's process so the final test performance remains an independent evaluation.
Identify the leak: The test performance influenced two model changes, so the test set guided calibration.
Add validation data: Use a separate validation set to guide the model changes while keeping the test set aside.
Reserve the test set: After calibration is complete, evaluate the resulting model on the test set for the final measurement.
The corrected process separates training for model building, validation for calibration, and testing for the independent final evaluation.
Key Takeaways
- Do not use test results to guide model calibration; reserve the test set for the independent final evaluation.
- Hold-out validation separates training, validation, and test roles.
- K-fold validation rotates the evaluation role across K equal partitions and averages the resulting scores.
- Repeated shuffled K-fold validation exposes the evaluation to different partition arrangements but requires training and evaluating P multiplied by K models when repeated P times.
- Choose the validation strategy with the dataset size, split stability, desired evaluation precision, and computational cost in mind.
Key Takeaways
- The test set must not guide model calibration because doing so removes its role as an independent final evaluation.
- Hold-out validation gives training, validation, and test data separate responsibilities.
- K-fold validation rotates evaluation across K partitions and averages the scores.
- Repeated shuffling creates new K-fold arrangements and improves exposure to partition variation, but increases computational cost.
- For P repeated runs with K folds per run, training and evaluating P multiplied by K models is required.