Concepts / Model Calibration

Model Calibration

The test set should not be used to guide model calibration.

  • Programming

The Protected Final Check

Model calibration means making choices or adjustments while developing a model. The central rule is that the test set should not guide those choices. If test results influence decisions about the model, the test set is no longer being held aside for an independent final evaluation.

assignassignhold asidebuildguide choicesevaluate once held asidemeasureAvailable dataTraining setBuild modelCalibrated modelValidation setGuide choicesFinal performanceestimateTest setFinal evaluation
How does using the test set to guide calibration cause information to leak into model decisions and make the final performance estimate unreliable?

Hold-Out Roles

The simplest validation protocol divides the available examples into three roles. Training data is used to build the model. Validation data is used while choosing or adjusting the model. Test data is held aside for the final performance measurement. This three-way separation is needed because a two-way division into training and test data is not enough when decisions are still being made.

dividedividehold asideuseuseuseAvailable examplesTraining dataBuild modelModel buildingValidation dataChoose or adjustModel calibrationTest dataFinal measureFinal evaluation
How are the available examples divided among training, validation, and test sets, and which set does each stage use?

A Three-Way Development Plan

Suppose a learner has a dataset and is still deciding between different model choices. Which role should each portion play?

Build: Use the training portion to build the candidate model.

Choose: Use the validation portion to compare or adjust candidate choices during development.

Protect: Do not use the test portion to guide those choices. Keep it aside for the final evaluation.

The validation portion supports calibration, while the test portion remains an independent final check.

Rotating K-Fold Evaluation

K-fold validation addresses some of the split-dependent variation of a single hold-out split. The data is divided into K equal partitions. One partition takes the validation role while the remaining partitions are used for training. The validation role then rotates so that each partition takes a turn. The resulting scores are averaged.

other partitions trainother partitions trainother partitions trainother partitions trainevaluate each runPartition 1Validation in run 1Remaining partitionsTraining in each runK scoresAveragePartition 2Validation in run 2Partition 3Validation in run 3Partition 4Validation in run 4
How does each partition take a turn as the validation set while the remaining partitions are used for training?

Four Partitions, Four Validation Turns

Imagine a K-fold process with four equal partitions.

Run 1: Partition 1 is used for validation, and the other three partitions are used for training.

Run 2: Partition 2 becomes the validation partition, while the other three are used for training.

Run 3: Partition 3 takes the validation role, and the remaining partitions are used for training.

Run 4: Partition 4 takes the validation role, and the remaining partitions are used for training.

Combine: Average the four resulting validation scores.

Every partition contributes a validation result, rather than relying on only one fixed validation portion.

Repeated Partitioning

When the dataset is especially small and a more precise evaluation is needed, K-fold validation can be repeated. Before each new run, the data is shuffled and divided into K partitions again. Each run therefore produces its own K-fold score. The final score is the average of the scores from all runs.

average fold scorescontinuecreate new foldsaverage fold scorescontributecontributeK-fold run 1First partition arrangementScore 1Average of run 1Shuffle dataRepartitionK-fold run 2New partition arrangementScore 2Average of run 2Final averageAcross runs
What changes when the data is reshuffled and repartitioned for another K-fold run, and how do the different splits contribute to the evaluation?

A single K-fold run explores one partition arrangement. Repeating the process exposes the evaluation to additional arrangements. This can help when performance measurements change substantially after different random shuffling and splitting choices.

The Cost of More Evidence

runrun K modelstrain and evaluateHold-out splitOne training and validationprocessOne set of runsLower costOne K-fold runK model runsP multiplied by KmodelsHigher costRepeated K-foldP runs, K models each
How does repeating K-fold validation multiply the number of training and evaluation runs compared with a single hold-out split?
ProtocolPartition behaviorEvaluation resultComputational implication
Simple hold-outOne fixed separation into training, validation, and test rolesOne validation measurement and a final test measurementSimplest protocol
K-fold validationThe validation role rotates across K equal partitionsAverage of the K resulting scoresRequires training and evaluating K models for one K-fold run
Repeated K-fold validationShuffle and repartition before each K-fold runAverage the scores from all runsRepeating the process P times requires training and evaluating P multiplied by K models

The benefit of repeated validation has a direct cost. One K-fold run requires training and evaluating models for each of its K partitions. If the process is repeated P times, the total number of trained and evaluated models is P multiplied by K. Repetition can provide more exposure to different splits, but it requires more computation.

Mistakes in Calibration

  • Using the test score to choose or adjust the model

    The test set has influenced the model decision, so it is no longer an independent final evaluation.

    Fix: Use a separate validation set to guide calibration and reserve the test set for the final measurement.

  • Assuming a single hold-out split is always reliable

    Small validation and test portions may not represent the broader data well.

    Fix: Treat instability across splitting rounds as a warning sign and consider K-fold validation.

  • Confusing K-fold validation with one fixed validation split

    That does not rotate the evaluation role across K partitions.

    Fix: Let each partition take a turn as validation, then average the resulting scores.

  • Ignoring the cost of repeated K-fold validation

    Each repetition performs another K-fold process.

    Fix: Remember that P repetitions with K partitions require training and evaluating P multiplied by K models.

Choose the Validation Protocol

MEDIUM

A dataset is very small, and performance measurements change noticeably after different random shuffling and splitting rounds. The test set has not yet been used. Which validation approach is the strongest candidate, and what trade-off should be explained?

Hints
  • Look for the method designed to respond to split-dependent variation.
  • Consider what happens when the K-fold process is repeated several times.

Selecting a More Stable Evaluation Process

A small dataset produces unstable measurements under different hold-out splits. What process should be considered?

Identify the warning sign: Large changes between random splitting rounds indicate split-dependent variation.

Use rotating evaluation: K-fold validation allows more portions of the data to take part in evaluation by rotating the validation role.

Consider repetition: If the dataset is especially small and more precise evaluation is needed, shuffle and repartition before additional K-fold runs.

State the trade-off: Additional runs expose evaluation to more partition arrangements but require training and evaluating more models.

K-fold validation, possibly repeated with shuffling, addresses split-dependent variation at increased computational cost.

Key Takeaways

  1. Do not use test results to guide model calibration; reserve the test set for an independent final evaluation.
  2. Simple hold-out validation separates training, validation, and test roles.
  3. K-fold validation rotates the validation role across equal partitions and averages the resulting scores.
  4. Repeated K-fold validation reshuffles and repartitions the data before each run, then averages the run scores.
  5. Repeated validation can reduce dependence on one partition arrangement, but it increases computation because P repetitions with K partitions require P multiplied by K model runs.

Key Takeaways

  • The test set must remain separate from calibration decisions.
  • Hold-out validation gives training, validation, and test data distinct roles.
  • K-fold validation rotates evaluation across K partitions and averages the scores.
  • Repeated K-fold validation uses new shuffles and partitions for additional runs.
  • More validation runs provide more split arrangements but increase computational cost.