True Risk
Approximation error measures the gap between the best hypothesis in a class and the true hypothesis.
Two Sources of Poor Performance
A model can perform poorly for two importantly different reasons. First, the hypothesis class may be too limited to represent the underlying relationship well. This is approximation error. Second, the learning algorithm may select a poor hypothesis from that class because it learned from only a finite sample. This is estimation error. The distinction matters because the two problems call for different remedies.
Approximation error asks what the best hypothesis in the class can achieve. Estimation error asks how far the learned hypothesis is from the true risk because it was learned from a finite sample.
Following Risk as Complexity Changes
Training risk is the risk measured on the sample used to fit the hypothesis. Empirical risk is the risk observed on that finite sample. As a model becomes more capable, its empirical risk can become small even when its performance on the broader data distribution is not equally good. The contrast between empirical risk and true risk is therefore more informative than training risk by itself.
The underfitting pattern is large empirical risk: the hypothesis performs poorly even on the sample used for learning. The overfitting pattern is different: empirical risk is small, but the gap between empirical risk and true risk is large. Validation risk helps reveal this gap because validation evaluates the learned hypothesis on data outside the training sample.
Reading the Learned Hypothesis
Two Models, Two Diagnoses
Imagine comparing two learned hypotheses using their training and validation performance. What can the patterns suggest?
Model A: Model A has high training risk and high validation risk. Because its empirical risk is already large, the pattern is consistent with underfitting.
Model B: Model B has low training risk but much higher validation risk. The large empirical-to-true-risk pattern is consistent with overfitting.
Interpretation: The observations help diagnose the learned hypothesis, but they do not directly reveal the approximation error of the whole hypothesis class.
High training risk points toward underfitting, while a large gap between training and validation performance points toward overfitting.
The training sample and the true data distribution answer different questions. Training performance describes how the learned hypothesis behaves on the finite sample used to fit it. True risk concerns performance on the true distribution. A learned hypothesis can therefore have low empirical risk and still have high true risk; the large gap is the pattern associated with overfitting.
Validation Beyond the Training Sample
Validation evaluates the learned hypothesis on a validation set rather than relying only on the sample used to fit it. This gives an estimate of the learned hypothesis's true risk. The source notes that the difference between distribution risk and validation risk can be bounded quite tightly using Theorem 11.1, which makes validation useful for assessing the learned hypothesis.
The Approximation Gap
Approximation error is about the gap between the true hypothesis and the best hypothesis available in a class. It concerns what the class can achieve in principle, not merely what one learning run produced. The learned hypothesis is only one member selected from that class, so its training risk cannot be read as a direct measurement of the class's approximation error.
Treating low training risk as proof that approximation error is low.
The learned hypothesis may fit the finite sample well while still having a large empirical-to-true-risk gap. That pattern is associated with overfitting.
Fix:
Use validation to assess the learned hypothesis and keep the class-level approximation question separate.Treating validation risk as a direct measurement of approximation error.
Validation assesses the learned hypothesis, which is only one member selected from the class.
Fix:
Interpret validation as evidence about the learned hypothesis's true risk, not as a direct estimate of the class's approximation gap.Assuming every poor result has the same remedy.
Poor performance can arise from a limited hypothesis class or from a poor hypothesis learned from finite data.
Fix:
Use training and validation patterns to identify which source of error is more plausible before choosing a remedy.
Choosing a Remedy
When empirical risk is large, underfitting is a plausible diagnosis, so the hypothesis class may be unable to represent the relationship well. When empirical risk is small but the empirical-to-true-risk gap is large, overfitting is a plausible diagnosis, so the learning result may be unreliable outside the training sample. These patterns do not prove the exact source of error, but they provide more informed guidance than training risk alone.
A learned hypothesis has high training risk and high validation risk. Another learned hypothesis has low training risk and substantially higher validation risk. Identify the more plausible pattern for each hypothesis and state whether the concern is primarily underfitting or overfitting.
Hints
- Start with the size of empirical or training risk.
- Then compare training performance with validation performance.
- A large gap between the two is diagnostically different from high risk on both.
Key Takeaways
- Approximation error is the gap between the true hypothesis and the best hypothesis in a class.
- Estimation error is the difference between empirical risk and true risk caused by learning from a finite sample.
- High empirical risk is associated with underfitting, while a large empirical-to-true-risk gap is associated with overfitting.
- Validation estimates the true risk of the learned hypothesis, but it does not directly measure the approximation error of the whole class.
- Training risk and validation evidence should be interpreted together before choosing a remedy.
Key Takeaways
- Approximation error concerns the limitation of the hypothesis class.
- Estimation error concerns the gap created by learning from a finite sample.
- Underfitting is associated with large empirical risk; overfitting is associated with a large empirical-to-true-risk gap.
- Validation helps estimate the true risk of a learned hypothesis outside its training sample.
- Neither training risk nor validation risk alone directly reveals approximation error.