Validation
Approximation error measures the gap between the best hypothesis in a class and the true hypothesis.
Two Sources of Poor Performance
When a learned predictor performs poorly, its error can have more than one cause. The hypothesis class may be too limited to represent the true relationship well. This is approximation error. Alternatively, the learning algorithm may have selected a poor hypothesis because it learned from only a finite sample. This is estimation error. These causes are different, so they suggest different remedies.
Approximation error concerns what the hypothesis class can achieve in principle. It measures the gap between the best hypothesis in that class and the true hypothesis. Estimation error concerns the hypothesis actually learned from a finite sample: it is the difference between empirical risk and true risk caused by learning from limited data.
Reading Risk Patterns
Training risk, also called empirical risk in this context, gives useful information about how well the learned hypothesis performs on the sample used for learning. A large empirical risk indicates underfitting: the learned hypothesis performs poorly even on the data used to fit it. By contrast, a large gap between empirical risk and true risk indicates overfitting: performance on the training sample is much better than performance beyond that sample.
Why Training Risk Is Not Enough
It is tempting to inspect only the training risk. That inspection can reveal underfitting when the risk is large, but it cannot reliably reveal approximation error. A learned hypothesis is only one member selected from the hypothesis class. Its training risk therefore does not tell us directly how well the best possible member of the class would perform against the true hypothesis.
Two Different Reasons for a Poor Result
Consider a learning task in which a learned predictor performs poorly. Which two explanations must be considered before choosing a remedy?
Check the hypothesis class: The class may be unable to represent the underlying relationship well. In that case, even its best member may remain far from the true hypothesis, indicating approximation error.
Check the learning result: The algorithm may have inferred a poor hypothesis from a finite sample. In that case, the learned hypothesis can differ substantially from what its class could achieve, indicating estimation error.
Use outside-sample evidence: A validation estimate examines the learned predictor on data outside the sample used to fit it. This helps assess the learned hypothesis and the gap between its empirical and true performance.
The same poor result can arise from a limited hypothesis class or from a poor finite-sample learning result. Validation helps diagnose the learned hypothesis, but it does not directly measure the approximation error of the whole class.
This distinction guides remedies. If underfitting is plausible because empirical risk is large, the hypothesis class or learning setup may be too limited. If overfitting is plausible because the empirical-to-true-risk gap is large, the problem is more consistent with the learned hypothesis failing to generalize beyond its finite training sample. The available risk patterns provide evidence, but they do not automatically identify every component of error.
Holding Out Validation Data
Validation uses some of the available training data as a validation set. The learning algorithm first receives training data and produces an output predictor. Some original training data is then set aside for validation rather than being used only as fitting evidence. The output predictor is evaluated on that validation set, and the observed success supplies information about its true risk.
The important change is the source of evidence. Instead of evaluating the predictor only on the same sample that helped produce it, validation evaluates it on held-out examples. This can provide a more accurate estimate of true risk than a broad, loose, or pessimistic bound that must cover every hypothesis and every possible data distribution.
Assessing a Learned Hypothesis
What do you think happens?
A predictor performs very well on the data used for fitting but poorly on the held-out validation set. Which error pattern does this most strongly suggest?
Reveal answer
Answer: Overfitting caused by a large empirical-to-true-risk gap
The predictor performs better on the fitting sample than outside that sample. That pattern is associated with a large gap between empirical risk and true risk. The result assesses the learned predictor; it does not directly reveal the approximation error of the entire hypothesis class.
Validation is therefore an assessment procedure for a learned hypothesis. It does not magically expose the true hypothesis or directly reveal how close the best member of the class is to it. Its value is narrower and practical: it provides evidence about how the selected predictor performs outside the fitting sample.
| Observation | What it suggests | What it does not prove |
|---|---|---|
| Large empirical risk | Underfitting may be present | The approximation error has been measured directly |
| Large empirical-to-true-risk gap | Overfitting may be present | The hypothesis class's best member is close to or far from the true hypothesis |
| Validation performance | Useful evidence about the learned predictor's true risk | A direct measurement of approximation error |
Risk observations answer different diagnostic questions.
Choosing Among Candidate Models
Validation is useful when selecting a model because candidate output predictors can be evaluated using the same validation procedure. The candidate with stronger validation success can be preferred: its validation result gives a more informative estimate of true risk than relying only on a broad, potentially pessimistic bound.
- Produce candidate output predictors from the learning process.
- Evaluate the candidates on the same validation procedure.
- Compare their validation success.
- Prefer the candidate whose validation result indicates stronger performance.
- Keep the interpretation focused on learned-predictor risk rather than treating validation as a direct approximation-error measurement.
Common Diagnostic Mistakes
Treating low training risk as proof that the model is good.
Low empirical risk can coexist with a large empirical-to-true-risk gap, which is associated with overfitting.
Fix:
Use validation performance to obtain evidence about how the learned predictor performs outside the fitting sample.Treating high training risk as a complete explanation of the problem.
Large empirical risk suggests underfitting, but it does not by itself identify the approximation error of the entire hypothesis class.
Fix:
Distinguish the selected hypothesis's training performance from what the best member of the class could achieve.Calling validation a direct measurement of approximation error.
Validation evaluates the learned predictor, which is only one member selected from the class.
Fix:
Interpret validation as evidence about the learned predictor's true risk and as a diagnostic aid.Ignoring estimation error when a model performs poorly.
The learning algorithm may instead have inferred a poor hypothesis from a finite sample.
Fix:
Consider both approximation error and estimation error before choosing a remedy.
Practice Check
A learning algorithm produces a predictor with large empirical risk. A validation procedure is then used, and the predictor also performs poorly on the validation set. Explain which error pattern is most consistent with these observations, what remains uncertain, and why validation is still useful.
Hints
- Start with the meaning of large empirical risk.
- Compare performance on the fitting sample with performance on held-out data.
- Separate conclusions about the learned predictor from conclusions about the whole hypothesis class.
- A strong answer identifies underfitting as the pattern most directly suggested by large empirical risk. If validation performance is also poor, the learned predictor is not performing well outside the fitting sample either. However, these observations do not directly measure approximation error, because approximation error concerns the best hypothesis available in the entire class. Validation remains useful because it provides evidence about the learned predictor's true risk and can support model comparison.
Key Takeaways
- Approximation error is the gap between the best hypothesis in a class and the true hypothesis.
- Estimation error is the difference between empirical risk and true risk caused by learning from a finite sample.
- Large empirical risk is associated with underfitting, while a large empirical-to-true-risk gap is associated with overfitting.
- Validation evaluates a learned predictor on held-out data to obtain a more informative estimate of its true risk.
- Validation supports model selection, but it does not directly measure the approximation error of the entire hypothesis class.