Concepts / Validation

Validation

Approximation error measures the gap between the best hypothesis in a class and the true hypothesis.

  • Programming

Two Sources of Poor Performance

When a learned predictor performs poorly, its error can have more than one cause. The hypothesis class may be too limited to represent the true relationship well. This is approximation error. Alternatively, the learning algorithm may have selected a poor hypothesis because it learned from only a finite sample. This is estimation error. These causes are different, so they suggest different remedies.

may includemay includePoor performanceApproximation errorBest class member versustrue hypothesisEstimation errorFinite-sample learningeffect
How can poor prediction performance be separated into the limitation of the hypothesis class and the effect of learning from finite data?

Approximation error concerns what the hypothesis class can achieve in principle. It measures the gap between the best hypothesis in that class and the true hypothesis. Estimation error concerns the hypothesis actually learned from a finite sample: it is the difference between empirical risk and true risk caused by learning from limited data.

Reading Risk Patterns

Training risk, also called empirical risk in this context, gives useful information about how well the learned hypothesis performs on the sample used for learning. A large empirical risk indicates underfitting: the learned hypothesis performs poorly even on the data used to fit it. By contrast, a large gap between empirical risk and true risk indicates overfitting: performance on the training sample is much better than performance beyond that sample.

patternpatternUnderfittingLarge empirical riskPoor trainingperformanceRisk remains large on thesampleOverfittingLargeempirical-to-true-risk gapPoor outside-sampleperformanceTraining success does notcarry over
What do training risk and the gap to true risk reveal about underfitting and overfitting?

Why Training Risk Is Not Enough

It is tempting to inspect only the training risk. That inspection can reveal underfitting when the risk is large, but it cannot reliably reveal approximation error. A learned hypothesis is only one member selected from the hypothesis class. Its training risk therefore does not tell us directly how well the best possible member of the class would perform against the true hypothesis.

Two Different Reasons for a Poor Result

Consider a learning task in which a learned predictor performs poorly. Which two explanations must be considered before choosing a remedy?

Check the hypothesis class: The class may be unable to represent the underlying relationship well. In that case, even its best member may remain far from the true hypothesis, indicating approximation error.

Check the learning result: The algorithm may have inferred a poor hypothesis from a finite sample. In that case, the learned hypothesis can differ substantially from what its class could achieve, indicating estimation error.

Use outside-sample evidence: A validation estimate examines the learned predictor on data outside the sample used to fit it. This helps assess the learned hypothesis and the gap between its empirical and true performance.

The same poor result can arise from a limited hypothesis class or from a poor finite-sample learning result. Validation helps diagnose the learned hypothesis, but it does not directly measure the approximation error of the whole class.

This distinction guides remedies. If underfitting is plausible because empirical risk is large, the hypothesis class or learning setup may be too limited. If overfitting is plausible because the empirical-to-true-risk gap is large, the problem is more consistent with the learned hypothesis failing to generalize beyond its finite training sample. The available risk patterns provide evidence, but they do not automatically identify every component of error.

Holding Out Validation Data

Validation uses some of the available training data as a validation set. The learning algorithm first receives training data and produces an output predictor. Some original training data is then set aside for validation rather than being used only as fitting evidence. The output predictor is evaluated on that validation set, and the observed success supplies information about its true risk.

part of datapart of datafitevaluate onmeasure performanceAvailable trainingdataFitting dataUsed by learning algorithmOutput predictorTrue-risk estimateValidation performanceValidation setHeld out for evaluation
How does data move when part of the available training data is held out for validation?

The important change is the source of evidence. Instead of evaluating the predictor only on the same sample that helped produce it, validation evaluates it on held-out examples. This can provide a more accurate estimate of true risk than a broad, loose, or pessimistic bound that must cover every hypothesis and every possible data distribution.

Assessing a Learned Hypothesis

What do you think happens?

A predictor performs very well on the data used for fitting but poorly on the held-out validation set. Which error pattern does this most strongly suggest?

  • Underfitting caused by large empirical risk
  • Overfitting caused by a large empirical-to-true-risk gap
  • Direct measurement of approximation error
Reveal answer

Answer: Overfitting caused by a large empirical-to-true-risk gap

The predictor performs better on the fitting sample than outside that sample. That pattern is associated with a large gap between empirical risk and true risk. The result assesses the learned predictor; it does not directly reveal the approximation error of the entire hypothesis class.

Validation is therefore an assessment procedure for a learned hypothesis. It does not magically expose the true hypothesis or directly reveal how close the best member of the class is to it. Its value is narrower and practical: it provides evidence about how the selected predictor performs outside the fitting sample.

ObservationWhat it suggestsWhat it does not prove
Large empirical riskUnderfitting may be presentThe approximation error has been measured directly
Large empirical-to-true-risk gapOverfitting may be presentThe hypothesis class's best member is close to or far from the true hypothesis
Validation performanceUseful evidence about the learned predictor's true riskA direct measurement of approximation error

Risk observations answer different diagnostic questions.

Choosing Among Candidate Models

Validation is useful when selecting a model because candidate output predictors can be evaluated using the same validation procedure. The candidate with stronger validation success can be preferred: its validation result gives a more informative estimate of true risk than relying only on a broad, potentially pessimistic bound.

  • Produce candidate output predictors from the learning process.
  • Evaluate the candidates on the same validation procedure.
  • Compare their validation success.
  • Prefer the candidate whose validation result indicates stronger performance.
  • Keep the interpretation focused on learned-predictor risk rather than treating validation as a direct approximation-error measurement.

Common Diagnostic Mistakes

  • Treating low training risk as proof that the model is good.

    Low empirical risk can coexist with a large empirical-to-true-risk gap, which is associated with overfitting.

    Fix: Use validation performance to obtain evidence about how the learned predictor performs outside the fitting sample.

  • Treating high training risk as a complete explanation of the problem.

    Large empirical risk suggests underfitting, but it does not by itself identify the approximation error of the entire hypothesis class.

    Fix: Distinguish the selected hypothesis's training performance from what the best member of the class could achieve.

  • Calling validation a direct measurement of approximation error.

    Validation evaluates the learned predictor, which is only one member selected from the class.

    Fix: Interpret validation as evidence about the learned predictor's true risk and as a diagnostic aid.

  • Ignoring estimation error when a model performs poorly.

    The learning algorithm may instead have inferred a poor hypothesis from a finite sample.

    Fix: Consider both approximation error and estimation error before choosing a remedy.

Practice Check

MEDIUM

A learning algorithm produces a predictor with large empirical risk. A validation procedure is then used, and the predictor also performs poorly on the validation set. Explain which error pattern is most consistent with these observations, what remains uncertain, and why validation is still useful.

Hints
  • Start with the meaning of large empirical risk.
  • Compare performance on the fitting sample with performance on held-out data.
  • Separate conclusions about the learned predictor from conclusions about the whole hypothesis class.
  1. A strong answer identifies underfitting as the pattern most directly suggested by large empirical risk. If validation performance is also poor, the learned predictor is not performing well outside the fitting sample either. However, these observations do not directly measure approximation error, because approximation error concerns the best hypothesis available in the entire class. Validation remains useful because it provides evidence about the learned predictor's true risk and can support model comparison.

Key Takeaways

  • Approximation error is the gap between the best hypothesis in a class and the true hypothesis.
  • Estimation error is the difference between empirical risk and true risk caused by learning from a finite sample.
  • Large empirical risk is associated with underfitting, while a large empirical-to-true-risk gap is associated with overfitting.
  • Validation evaluates a learned predictor on held-out data to obtain a more informative estimate of its true risk.
  • Validation supports model selection, but it does not directly measure the approximation error of the entire hypothesis class.