Concepts / True Risk Estimation

True Risk Estimation

Validation uses some training data as a validation set.

  • Programming

The Question After Training

A learning algorithm produces an output predictor from training data. The next question is not merely whether the algorithm produced a predictor, but how well that predictor will perform in terms of its true risk. Evidence from the fitting process alone may not provide as informative an estimate as we want. Validation supplies a separate procedure for gathering better evidence about the predictor.

What do you think happens?

A predictor was produced from the available training data. What additional action can provide more information about its true risk?

  • Evaluate the output predictor on a validation set formed from some set-aside training data
  • Rely only on a broad bound that covers every hypothesis in the class
  • Choose a model without evaluating its output predictor
Reveal answer

Answer: Evaluate the output predictor on a validation set formed from some set-aside training data

Validation sets aside some of the training data and evaluates the output predictor on that set. The observed success provides information for estimating true risk.

Splitting the Available Data

Validation begins with the training data already available to the learning task. Some of that data is set aside as a validation set. The remaining portion is used by the learning algorithm to produce an output predictor. The predictor is then evaluated on the set-aside portion rather than relying only on the evidence used to produce it.

divideset asideproduceevaluate on validation setsupplies examplesTraining dataavailable examplesFit portionused by learning algorithmOutput predictorproduced from fit portionValidation successevidence for true riskValidation setset aside for evaluation
How is the available training data divided so that one portion produces the predictor and another portion evaluates it?

A Predictor's Validation Result

One Predictor, Two Roles for Data

A learning task has available training data. The task must estimate how well its output predictor will perform in terms of true risk.

Produce the predictor: The learning algorithm receives the portion of the training data used for fitting and produces an output predictor.

Set aside validation data: Some of the original training data is used as a validation set rather than as the fitting evidence for that predictor.

Evaluate success: The output predictor is evaluated on the validation set. This observed success supplies information about the predictor's true risk.

Interpret the estimate: The validation result is used as a more accurate estimate of true risk than relying only on a broad, potentially pessimistic bound.

Validation creates a separate evaluation step: one portion helps produce the predictor, and the set-aside portion helps estimate how the predictor will perform.

This example does not claim that validation reveals the exact true risk. It shows the role of validation: the predictor's observed success on the validation set becomes evidence used to estimate true risk. The value comes from evaluating the algorithm's output on data that was set aside for that purpose.

Why Training Evidence Can Mislead

A broad theoretical bound can relate true risk to empirical risk for every hypothesis in a hypothesis class. However, such a bound must cover all hypotheses and all possible data distributions. That broad coverage can make the bound loose and pessimistic. Validation changes the source of the estimate by examining the particular output predictor on the validation set.

produceevaluatesupplies evidenceFitting evidenceused to produce predictorOutput predictorproduced by algorithmValidation evidenceset-aside evaluation dataOutput predictorevaluated on validation setTrue-risk estimatevalidation success
What is the difference between evidence used to produce a predictor and evidence obtained by evaluating it on a set-aside validation set?

When reasoning about a predictor's true risk, identify where the estimate comes from. A broad bound provides general coverage, while validation evaluates the particular output predictor on a set-aside validation set. The source describes the latter as providing a more accurate estimate.

Choosing Between Candidate Predictors

Validation is useful when selecting a model because competing output predictors can be evaluated using the same validation procedure. The candidate with stronger validation success can be preferred: its validation result gives a better estimate of its true risk than relying only on a broad, potentially pessimistic bound.

evaluateevaluatecompare resultsCandidate Avalidation successSame validationprocedurecommon evaluationPreferred candidatestronger validation successCandidate Bvalidation success
How does validation performance guide a choice between competing output predictors?

Suppose a learning task produces two candidate output predictors. If both are evaluated through the same validation procedure, the candidate with stronger validation success can be preferred. The comparison is based on validation evidence about the candidates' true risk rather than only on a broad bound.

Common Reasoning Mistakes

  • Treating the fitting evidence and validation evidence as having the same role

    Validation is specifically a separate procedure in which some training data is set aside and the output predictor is evaluated on it.

    Fix: Distinguish the portion used by the learning algorithm from the portion used to evaluate the resulting predictor.

  • Treating a broad bound as the only available estimate

    Bounds that cover every hypothesis and possible data distribution can be loose and pessimistic.

    Fix: Use validation to obtain information from the particular output predictor's success on the validation set.

  • Choosing a model without comparing validation performance

    Validation can provide a better estimate of the candidates' true risks and is useful for selecting a model.

    Fix: Evaluate the competing candidates using the same validation procedure and compare their validation success.

Check Your Understanding

MEDIUM

A learning algorithm produces an output predictor after receiving part of the available training data. Explain what the set-aside validation data contributes, why its result can be more informative than a broad bound, and how the result could help choose between two candidate predictors.

Hints
  • Start by naming the role of the validation set.
  • Explain what is evaluated on that set.
  • Connect stronger validation success with model selection.
  1. Validation uses some available training data as a validation set. The learning algorithm produces an output predictor from the fitting portion, and that predictor is evaluated on the set-aside portion. The observed validation success provides information for estimating true risk and can be more accurate than broad, loose, or pessimistic bounds. Because competing predictors can be evaluated through the same procedure, validation also supports model selection.

Key Takeaways

  • Validation sets aside some training data for evaluating the output predictor.
  • The validation result provides information for estimating the predictor's true risk.
  • Broad bounds can be loose and pessimistic because they must cover every hypothesis and possible data distribution.
  • Validation is useful for comparing candidate predictors and selecting a model.