Concepts / Information Leakage

Information Leakage

The test set should not be used to guide model calibration.

  • Programming

The Test Set Must Stay Independent

Information leakage occurs when information from the test set influences model calibration. Calibration includes choosing or adjusting the model. A test set is meant to provide an independent final evaluation, so using its results to guide decisions changes the role of that data.

measureinfluenceschanges reliabilityTest setfinal evaluation dataTest scoreobserved performanceModel decisionchoose or adjustFinal estimateno longer independent
How does information from the test set flow backward into model decisions, and why does that make the final performance estimate unreliable?

Choosing Between Two Models

A learner builds two candidate models and uses performance measurements to decide which one to keep. Which dataset should guide that decision?

Separate the roles: Keep training data for building the model, use a separate validation set to guide the choice or adjustment, and reserve the test set for the final evaluation.

Make the calibration decision: Compare the candidate models using validation results rather than test results. The validation results are allowed to influence the model decision.

Evaluate once at the end: After the model has been chosen or adjusted, use the held-aside test set for the independent final evaluation.

The validation set guides calibration; the test set is reserved for the final evaluation.

Three Roles in Hold-Out Validation

Simple hold-out validation assigns separate roles to three portions of the data. The training portion is used to build the model. The validation portion is used while choosing or adjusting the model. The test portion is held aside for the final evaluation. This separation prevents calibration decisions from being guided by the data intended for the final check.

assignassignassignbuildguideevaluate with test dataprovide final dataDatasetseparated into rolesTraining databuild the modelCalibrated modelafter validation decisionsFinal evaluationtest performanceValidation datachoose or adjustTest dataheld aside
What data is used at each stage, and which datasets are allowed to influence model fitting, calibration, and final evaluation?

Hold-out validation is the simplest protocol, but it can be unreliable when only a small amount of data is available. The validation and test portions may contain too few samples to represent the broader data well. Instability is a warning sign: different random shuffling and splitting rounds produce very different performance measurements.

K-Fold Rotation

K-fold validation addresses split-dependent variation by dividing the data into K equal partitions. One partition takes the evaluation role while the remaining partitions are used for training. The evaluation role then rotates, so each partition serves as the evaluation portion once. The resulting scores are averaged.

rotaterotatecontinueaveragePartition 1validationAverage scorescores from all roundsPartition 2validationPartition 3validationPartition Kvalidation
How does each partition take turns serving as the validation set while the remaining partitions are used for training?

A Four-Partition Rotation

A dataset is divided into four equal partitions. How does one K-fold validation run use those partitions?

Round 1: Partition 1 serves as the evaluation portion, while the other three partitions are used for training.

Round 2: Partition 2 serves as the evaluation portion, while the other three partitions are used for training.

Continue the rotation: Partitions 3 and 4 each take the evaluation role in turn, with the remaining partitions used for training in each round.

Combine the results: The scores from the four rounds are averaged to produce the K-fold score for this run.

Every partition contributes once as the evaluation portion, and the run produces an average score.

Use K-fold validation when a single hold-out split may be too dependent on which samples happened to land in the validation portion. It allows more portions of the data to take part in evaluation, while still rotating the training and evaluation roles.

Repeated Shuffling and Cost

When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each run, the data is shuffled and divided into K partitions again. Each run therefore uses a new partition arrangement, produces its own K-fold score, and contributes that score to the final average.

train and evaluaterepeattrain and evaluateincludeincludeShuffle 1K partitionsK-fold score 1K evaluationsFinal averagescores from all runsShuffle 2new K partitionsK-fold score 2K evaluations
How do different shuffles create new K-fold partitions, and why does repeating the process require additional model training runs?

Repeated shuffling increases exposure to different partition arrangements, but it also increases computational cost. If the process is repeated P times and each run uses K folds, training and evaluating more models is required: P multiplied by K models in total.

Validation Versus Final Testing

AspectValidation performanceTest performance
RoleGuide model choice or adjustmentProvide the final evaluation
May influence calibration?YesNo
Data statusUsed during the decision processHeld aside for an independent final evaluation
If used incorrectlyThe calibration process may be affectedThe final estimate is no longer independent

The difference is about direction of influence. Training data influences model building. Validation data can influence calibration. Test data should receive the final model without influencing the decisions that produced it. Confusing these directions is the central information-leakage mistake.

Mistakes That Cause Leakage

  • Using the test score to choose between candidate models

    The test results have influenced model calibration, so the test set is no longer reserved for an independent final evaluation.

    Fix: Use a separate validation set to guide the choice, then use the test set for the final evaluation.

  • Treating a training-and-test split as sufficient while still adjusting the model

    The test results are guiding repeated calibration decisions.

    Fix: Introduce a validation role between model building and final testing.

  • Assuming one hold-out split is always reliable

    Small validation and test portions may not represent the broader data well, and instability signals split-dependent variation.

    Fix: Consider K-fold validation, or repeated shuffled K-fold validation when the dataset is especially small and evaluation precision is important.

  • Repeating K-fold validation without accounting for computational cost

    Each repeated run contains K training and evaluation rounds.

    Fix: For P repeated runs and K folds per run, account for training and evaluating P multiplied by K models.

Practice the Data Roles

MEDIUM

A learner first builds a model with training data, then changes the model twice after looking at test performance. The learner reports the last test performance as the final result. Identify the leakage, name the missing data role, and describe how the process should be reorganized.

Hints
  • Ask which dataset influenced the model changes.
  • The final evaluation requires data that was not used to guide calibration.
  • Name the separate role that should guide choosing or adjusting the model.

Practice Solution

Reorganize the learner's process so the final test performance remains an independent evaluation.

Identify the leak: The test performance influenced two model changes, so the test set guided calibration.

Add validation data: Use a separate validation set to guide the model changes while keeping the test set aside.

Reserve the test set: After calibration is complete, evaluate the resulting model on the test set for the final measurement.

The corrected process separates training for model building, validation for calibration, and testing for the independent final evaluation.

Key Takeaways

  1. Do not use test results to guide model calibration; reserve the test set for the independent final evaluation.
  2. Hold-out validation separates training, validation, and test roles.
  3. K-fold validation rotates the evaluation role across K equal partitions and averages the resulting scores.
  4. Repeated shuffled K-fold validation exposes the evaluation to different partition arrangements but requires training and evaluating P multiplied by K models when repeated P times.
  5. Choose the validation strategy with the dataset size, split stability, desired evaluation precision, and computational cost in mind.

Key Takeaways

  • The test set must not guide model calibration because doing so removes its role as an independent final evaluation.
  • Hold-out validation gives training, validation, and test data separate responsibilities.
  • K-fold validation rotates evaluation across K partitions and averages the scores.
  • Repeated shuffling creates new K-fold arrangements and improves exposure to partition variation, but increases computational cost.
  • For P repeated runs with K folds per run, training and evaluating P multiplied by K models is required.