Empirical Risk
Approximation error measures the gap between the best hypothesis in a class and the true hypothesis.
The Evidence a Learner Can See
A learning algorithm receives a training set sampled from an unknown distribution and labeled by an unknown target function. It must choose a predictor without directly inspecting the entire source distribution or the target function. The error measured on the available training sample is observable to the learner; the error over the unknown distribution and target function is not directly available.
Empirical risk is the error a predictor incurs on the training sample. It gives the learner a concrete basis for comparing predictors, but it is not automatically the same as the predictor's true risk outside that sample.
Two Sources of Error
A poor learned result can have more than one explanation. The hypothesis class may be unable to represent the underlying relationship well. Alternatively, the class may be capable of doing well, but the learning process may select a poor hypothesis because it learned from only a finite sample. These possibilities are described by approximation error and estimation error.
Approximation error is the gap between the true hypothesis and the best hypothesis available in a chosen hypothesis class. It concerns what the class can achieve in principle.
Estimation error is the difference between empirical risk and true risk caused by learning from a finite sample. It concerns the gap created when the learner infers a particular hypothesis from limited evidence.
How ERM Chooses a Predictor
Empirical Risk Minimization, or ERM, selects a predictor by minimizing its error on the training sample.
The learner cannot directly minimize true risk because it cannot directly inspect the entire unknown distribution or the unknown target function. Instead, it uses the training set as evidence. Candidate predictors are compared according to the errors they incur on that sample, and ERM favors the predictor with the smallest measured error.
Comparing Two Candidate Predictors
A training sample is used to compare Predictor A and Predictor B. Predictor A makes fewer errors on the training sample than Predictor B. Which predictor does ERM favor?
Measure: Evaluate the errors made by both predictors on the training sample.
Compare: Place the measured empirical risks side by side.
Select: ERM favors the predictor with the smaller measured training error.
ERM favors Predictor A. This conclusion concerns empirical risk on the training sample; it does not by itself establish which predictor has lower true risk.
What do you think happens?
Predictor A has a smaller training error than Predictor B. Which predictor does ERM favor?
Reveal answer
Answer: Predictor A
ERM uses the error measured on the training sample and favors the predictor with the minimum measured error. This choice does not by itself prove that Predictor A has lower true risk.
Validation Beyond the Training Sample
Validation evaluates a learned hypothesis on data separate from the training sample. Because the hypothesis was not fitted using those validation examples, validation provides evidence about how the learned hypothesis performs outside the sample used to fit it. The source describes validation as a way to estimate the true risk of the learned hypothesis, with the difference between distribution risk and validation risk capable of being bounded quite tightly using Theorem 11.1.
Reading Risk Patterns
Training and validation or true risk together provide more information than training risk alone. Large empirical risk indicates underfitting: the predictor performs poorly even on the training sample. A large gap between empirical risk and true risk indicates overfitting: the predictor appears better on the training sample than it does outside that sample.
| Observed pattern | Plausible issue | Interpretation |
|---|---|---|
| Large empirical risk | Underfitting | The predictor performs poorly on the training sample. |
| Small empirical risk with a large empirical-to-true-risk gap | Overfitting | The predictor performs much better on the training sample than outside it. |
| Training risk by itself | Insufficient diagnosis | It does not necessarily reveal approximation error. |
Choosing a Remedy
When a model performs poorly, first ask which error source is most plausible. If the hypothesis class is too limited to represent the underlying relationship well, the concern is approximation error and a more expressive or otherwise different class may be worth considering. If the learned hypothesis performs well on training data but poorly outside it, the concern is estimation error or overfitting; more data or stronger regularization may be relevant remedies. These choices are diagnostic directions, not guarantees, because validation does not directly measure approximation error.
Common Diagnostic Mistakes
Treating training risk as a direct measurement of approximation error.
Training risk concerns the selected predictor on one finite sample. Approximation error concerns the gap between the true hypothesis and the best member of the entire class.
Fix:
Use training and validation or true-risk evidence to assess the learned hypothesis, while remembering that approximation error remains harder to estimate directly.Assuming validation directly reveals approximation error.
Validation evaluates the particular learned hypothesis, not necessarily the best possible member of the class.
Fix:
Use validation to assess the learned hypothesis outside the training sample, not as a complete measurement of approximation error.Assuming ERM minimizes true risk.
The learner cannot directly inspect the unknown distribution or target function.
Fix:
State that ERM minimizes measured error on the training sample.Calling every poor result underfitting.
Poor performance can arise from a limited class or from learning a poor hypothesis from a finite sample.
Fix:
Compare empirical risk with outside-sample performance to investigate approximation and estimation concerns separately.
Check Your Understanding
A learning algorithm compares two predictors on a training sample. Predictor R has the smaller training error, but its performance on a separate validation set is much worse than its training performance. Explain which predictor ERM selects, what pattern the validation result suggests, and why the result still does not directly measure approximation error.
Hints
- ERM selects using error measured on the training sample.
- Compare empirical risk with performance outside the training sample.
- Approximation error concerns the best member of the hypothesis class, not only the selected predictor.
Answering the Diagnostic Question
Predictor R has the smallest training error, but its validation error is much larger.
Selection: ERM selects Predictor R because it has the smallest measured error on the training sample.
Pattern: The large difference between training and validation performance suggests a large empirical-to-true-risk gap and therefore an overfitting or estimation concern.
Limitation: The validation result concerns the learned Predictor R. It does not directly tell us how well the best member of the whole hypothesis class could perform.
ERM selection, overfitting diagnosis, and approximation-error diagnosis are three related but distinct conclusions.
Key Takeaways
- Empirical risk is the error measured on a finite training sample and is observable to the learner.
- ERM selects a predictor by favoring the smallest training-sample error because true risk is not directly available.
- Approximation error concerns the gap between the true hypothesis and the best member of a hypothesis class.
- Estimation error concerns the gap between empirical risk and true risk caused by learning from a finite sample.
- Large empirical risk suggests underfitting, while a large empirical-to-true-risk gap suggests overfitting.
- Validation helps assess the learned hypothesis outside the training sample, but it does not directly measure approximation error.
Key Takeaways
- Empirical risk measures observed performance on a finite training sample, not automatically performance over the unknown distribution.
- ERM uses training error because the learner cannot directly access true risk.
- Approximation error describes a limitation of the hypothesis class, while estimation error comes from learning from finite data.
- Validation provides evidence about the learned hypothesis outside the training sample but does not directly reveal approximation error.
- Training and validation patterns help distinguish likely underfitting from likely overfitting and guide more informed remedies.