Concepts / Parameter Tuning

Parameter Tuning

k-Fold Cross Validation helps estimate true error without setting aside a permanently unused validation set.

  • Programming

The Scarce-Data Problem

A validation procedure usually needs some data to be held aside so that a model can be checked. When the available data is scarce, permanently reserving a separate validation set means that those examples cannot contribute to training. k-Fold Cross Validation addresses this problem by repeatedly using different parts of the original training set for validation while using the remaining parts for training.

replacereuse examplesPermanentvalidation setheld asideRotating validationfoldsdifferent examples eachfoldTraining setremaining examplesRotating trainingfoldsremaining examples eachfold
How does k-Fold Cross Validation let the same scarce examples contribute to both model training and error estimation?

One Complete Cross-Validation Pass

Choose a value for k and divide the original training set into k folds. The procedure then repeats k times. During one repetition, one fold serves as the error-estimation fold, while all the other folds are used for training. On the next repetition, a different fold serves as the error-estimation fold. Each fold serves as the error-estimation fold exactly once.

rotate held-out foldcontinueFold 1train on 2,3,...,k;estimate on 1Fold 2train on 1,3,...,k;estimate on 2Fold ktrain on 1,2,...,k-1;estimate on k
How do the training and validation sets change across each fold, and how is every example used for both training and error estimation?

The important state change from one fold to the next is not a change to the original data. It is a change in which fold is used to estimate error and which folds are used to train. Across the complete pass, every example belongs to a training set during some folds and to an error-estimation fold once.

From Fold Errors to True Error

Each fold produces an error measurement for a model trained on the other folds. The average of the k fold errors is the cross-validation estimate of the model's true error. This estimate uses the repeated error measurements from the different held-out parts instead of relying on one permanently unused validation set.

combinecombinecombineestimateFold 1 errorheld-out fold 1Average fold errorcross-validation estimateTrue errorestimatedFold 2 errorheld-out fold 2Fold k errorheld-out fold k
How do the individual validation errors from the folds combine into an average that estimates the model's error on unseen data?

Comparing Two Parameter Values

A learner evaluates two possible parameter values using three folds. Parameter A produces fold errors of 2, 4, and 3. Parameter B produces fold errors of 5, 2, and 4. Which parameter has the lower cross-validation error?

Collect the fold errors: For each parameter, retain the error produced when each of the three folds serves as the error-estimation fold.

Average parameter A: The three errors for parameter A average to 3.

Average parameter B: The three errors for parameter B average to 11 divided by 3, which is approximately 3.67.

Compare the estimates: The lower average is the lower estimated true error for this comparison.

Parameter A is selected because its cross-validation error is lower than parameter B's.

Selecting a Parameter

k-Fold Cross Validation can be used for model selection, also called parameter tuning. Give the procedure a training set, a set of possible parameter values, a choice of k, and a learning algorithm that accepts both a training set and a parameter. Run cross-validation for each possible parameter value, compare the resulting cross-validation errors, and choose the parameter with the lowest estimated error.

lower estimatecompareParameter Aaverage error: 3Parameter AselectedParameter Baverage error: 3.67
How are errors for different parameter values compared across folds to choose the parameter with the lowest estimated error?
Candidate parameterFold 1 errorFold 2 errorFold 3 errorCross-validation error
Parameter A2433
Parameter B524Approximately 3.67

Illustrative comparison of candidate parameter values using three folds.

Retraining After Selection

Cross-validation is used to choose the parameter, not to produce the final model that is ready for use. After the best parameter has been chosen, retrain the learning algorithm with that parameter on the entire original training set. This final training step allows the selected model to use all of the original training examples.

evaluatechoose lowest estimatepass parameterproduceTraining set andcandidatespossible parameter valuesk-Fold CrossValidationestimate each candidateBest parameterlowest estimated errorRetrain algorithmentire original trainingsetFinal modelselected parameter
What happens to the selected parameter after cross-validation, and how does the final model get retrained on the entire training set?

Leave-One-Out Cross Validation

Leave-One-Out Cross Validation is the special case of k-Fold Cross Validation in which k equals m, where m is the number of examples. In this arrangement, the fold structure makes each example the single error-estimation case in turn, while the other examples are used for training.

all other examples trainall other examples trainall other examples trainExample 1validation caseRemaining examplestraining casesExample 2validation caseExample mvalidation case
What changes in the fold structure when k equals the number of examples, and how does each example become the single validation case in turn?

The defining feature is the value of k: it equals the number of examples. The validation fold therefore contains one example at a time, and the held-out example changes until every example has served as the error-estimation case.

Common Tuning Mistakes

  • Keeping one permanently unused validation set when the training data is scarce.

    The held-aside examples cannot help train the model, even though the available data is limited.

    Fix: Use k-Fold Cross Validation so different parts of the original training set take turns serving as the error-estimation fold.

  • Using only one fold as the validation fold.

    The procedure requires each fold to serve as the error-estimation fold exactly once.

    Fix: Complete all k repetitions and combine the k fold errors.

  • Choosing a parameter from one fold error instead of the cross-validation error.

    A single fold does not provide the full cross-validation estimate described by the procedure.

    Fix: Average the errors from all folds for each candidate parameter, then compare those averages.

  • Failing to retrain after selecting the parameter.

    The stated model-selection procedure ends with retraining on the entire original training set.

    Fix: Use the selected parameter with the learning algorithm on all original training examples.

Practice the Procedure

MEDIUM

A training set is divided into four folds. A candidate parameter produces four fold errors: 6, 3, 5, and 2. Describe what each error represents, explain what should be done with the four errors, and state what happens after this candidate is compared with the other possible parameter values.

Hints
  • Identify which fold is used for error estimation during each repetition.
  • The other folds are used for training during that repetition.
  • The four errors are combined into the cross-validation estimate for this candidate.
  • After comparison, retrain with the selected parameter on the entire original training set.

What do you think happens?

Suppose k equals the number of examples. What does each error-estimation fold contain?

  • All examples
  • Half of the examples
  • Exactly one example
  • No examples
Reveal answer

Answer: Exactly one example

When k equals m, the procedure is Leave-One-Out Cross Validation. Each example serves as the single error-estimation case in turn, while the other examples are used for training.

Key Takeaways

  1. k-Fold Cross Validation estimates true error without setting aside a permanently unused validation set.
  2. Each fold is used for error estimation exactly once, while the other folds are used for training.
  3. The average of the fold errors is the cross-validation estimate of true error.
  4. Leave-One-Out Cross Validation is the special case where k equals the number of examples.
  5. For parameter tuning, compare candidate parameters by their cross-validation errors, select the lowest estimate, and retrain on the entire original training set.

Key Takeaways

  • k-Fold Cross Validation is useful when data is scarce because examples rotate between training and error estimation rather than remaining permanently unused.
  • Across k folds, each fold serves as the error-estimation fold once and the remaining folds serve as training data.
  • The average fold error estimates the model's true error.
  • Leave-One-Out Cross Validation occurs when k equals the number of examples.
  • Parameter tuning uses these estimates to select a parameter before retraining on the entire original training set.