Hypothesis Selection
Overfitting is a mismatch between excellent training performance and poor true-distribution performance.
The Training-Performance Trap
A hypothesis can classify every example in the observed training set correctly and still perform poorly on examples from the true data distribution. This is the central warning of overfitting: excellent performance on the observed sample does not necessarily indicate excellent performance in the true world.
Two Ways to Evaluate a Hypothesis
Hypothesis selection involves comparing how a hypothesis performs on the training data with how it performs on the true data distribution. Performance measured on the training data is represented by empirical error. Performance measured on the true data distribution is represented by true error. Overfitting occurs when these two views do not agree: the hypothesis has excellent training performance but poor true-distribution performance.
The diagram separates the two evaluations. The same hypothesis h_S has zero empirical error on the training data but true error equal to 1/2 on the true distribution. The result is not contradictory: the evaluation set has changed.
Why ERM Can Overfit
An empirical risk minimization algorithm selects a hypothesis because it minimizes empirical error. That objective evaluates performance on the training data. Therefore, ERM may select a hypothesis that fits the training data extremely well even when that hypothesis performs poorly on the true data distribution. The selection is successful according to the empirical criterion, but the selected hypothesis can still be overfitting.
The important distinction is between the selection rule and the broader goal. ERM uses empirical error to choose a hypothesis. Overfitting appears when the hypothesis favored by that training-based rule does not perform well on the true data distribution.
A Worked Error Pattern
Reading the h_S Example
Interpret the example in which L_S(h_S) = 0 and L_D(h_S) = 1/2.
Read the empirical error: L_S(h_S) = 0 means that h_S has zero error when evaluated on the training data.
Read the true error: L_D(h_S) = 1/2 means that h_S has nonzero error when evaluated on the true data distribution.
Compare the evaluations: The hypothesis is perfect under the training-data evaluation but performs poorly under the true-distribution evaluation.
Name the pattern: This mismatch is overfitting.
Zero empirical error does not imply zero true error. The example demonstrates why training success alone is not enough.
In this pattern, h_S is the overfitting hypothesis because its training performance is better than its true-distribution performance. The defining signal is not merely that one error value is large. It is the mismatch between excellent performance on the observed training data and poor performance on the true data distribution.
Mistakes in Interpretation
Concluding that zero empirical error means zero true error.
The source example has L_S(h_S) = 0 but L_D(h_S) = 1/2.
Fix:
Evaluate training performance and true-distribution performance separately.Assuming ERM always selects the hypothesis with the best true-distribution performance.
ERM may select an overfitting hypothesis because it minimizes empirical error.
Fix:
Remember that the ERM selection criterion is empirical error.Using the word overfitting for any hypothesis with a nonzero error.
Overfitting specifically describes a mismatch between excellent training performance and poor true-distribution performance.
Fix:
Look for the contrast between the two evaluation settings.
Check the Pattern
A hypothesis has excellent performance on the training data but poor performance on the true data distribution. What concept does this pattern demonstrate, and why might ERM select this hypothesis?
Hints
- Compare the two evaluation settings rather than looking only at the training result.
- Recall which error ERM minimizes.
Key Takeaways
- Overfitting is a mismatch between excellent training performance and poor true-distribution performance.
- Zero empirical error does not guarantee zero true error.
- ERM may select an overfitting hypothesis because it minimizes empirical error.
- The example L_S(h_S) = 0 and L_D(h_S) = 1/2 shows why training success alone is not enough.
- To identify overfitting, compare a hypothesis's performance on the training data with its performance on the true data distribution.
Key Takeaways
- Overfitting means excellent performance on the training data combined with poor performance on the true data distribution.
- ERM can choose an overfitting hypothesis because it minimizes empirical error.
- The values L_S(h_S) = 0 and L_D(h_S) = 1/2 show that zero empirical error can coexist with nonzero true error.
- The reliable way to identify overfitting is to compare training performance with true-distribution performance.