True Error and Empirical Error
Overfitting is a mismatch between excellent training performance and poor true-distribution performance.
The Training-World Trap
A hypothesis can perform perfectly on the training set and still perform poorly on examples from the true data distribution. This is the central warning of overfitting: success on the observed sample does not necessarily mean success in the true world.
Overfitting is a mismatch between excellent training performance and poor true-distribution performance.
Two Ways to Measure Performance
Empirical error describes how a hypothesis performs on the observed training set. True error describes how that hypothesis performs on the true data distribution. The two measures can disagree because excellent performance on the training sample does not guarantee excellent performance in the true world.
Zero empirical error means that the hypothesis has no observed error on the training set. It does not mean that the hypothesis has zero true error.
A Perfect Training Score
The h_S Example
Interpret the case in which the empirical error of h_S is 0 while its true error is 1/2.
Read the empirical error: The value 0 means that h_S has no error on the training set.
Read the true error: The value 1/2 means that h_S has nonzero error on the true data distribution.
Compare the two results: The hypothesis is perfect on the observed training data but performs poorly on the true distribution.
Name the pattern: This mismatch is overfitting.
Training success alone is not enough to establish good true-distribution performance.
Why ERM Can Overfit
An empirical risk minimization algorithm selects a hypothesis because it minimizes empirical error. Its selection is therefore based on performance measured on the training data. If a hypothesis has the lowest observed training error but poor performance on the true data distribution, ERM may select that overfitting hypothesis.
ERM may choose an overfitting hypothesis not because its true error is lowest, but because its empirical error is lowest.
Following the Mismatch
To recognize overfitting, compare the hypothesis's performance on the training data with its performance on the true data distribution. The overfitting pattern appears when training performance is excellent while true-distribution performance is poor.
Whenever a training result looks excellent, ask a second question: how does the hypothesis perform on the true data distribution? This comparison is necessary for detecting the mismatch that defines overfitting.
Mistakes in Error Interpretation
Treating zero empirical error as zero true error.
The zero applies to the training set, while true error concerns the true data distribution.
Fix:
State precisely which error is zero and check the other measure separately.Assuming the hypothesis selected by ERM must have the lowest true error.
ERM minimizes empirical error, and the source example shows that empirical success can coexist with nonzero true error.
Fix:
Describe ERM's selection criterion as empirical error and avoid equating it automatically with true-distribution performance.Ignoring the difference between the observed sample and the true data distribution.
Overfitting is precisely the mismatch between these two performance settings.
Fix:
Name the evaluation setting whenever you report a performance result.
Check Your Interpretation
A hypothesis has empirical error 0 and true error 1/2. What does this tell you about its training performance, its true-distribution performance, and whether the pattern is overfitting?
Hints
- Interpret empirical error using the training set.
- Interpret true error using the true data distribution.
- Compare the two results before naming the pattern.
What do you think happens?
If ERM selects a hypothesis with the lowest empirical error, can that selected hypothesis still be overfitting?
Reveal answer
Answer: Yes
ERM minimizes empirical error, but the selected hypothesis can still have poor performance on the true data distribution.
Key Takeaways
- Empirical error concerns performance on the observed training set.
- True error concerns performance on the true data distribution.
- Overfitting is the mismatch between excellent training performance and poor true-distribution performance.
- A hypothesis can have zero empirical error and nonzero true error, as in the example with true error 1/2.
- ERM may select an overfitting hypothesis because it minimizes empirical error.
Key Takeaways
- Empirical error and true error measure performance in different settings.
- Zero empirical error does not guarantee zero true error.
- Overfitting occurs when training performance is excellent but true-distribution performance is poor.
- ERM can select an overfitting hypothesis because its selection is based on empirical error.
- The pattern L_S(h_S) = 0 and L_D(h_S) = 1/2 demonstrates why training success alone is not enough.