Concepts / True Error and Empirical Error

True Error and Empirical Error

Overfitting is a mismatch between excellent training performance and poor true-distribution performance.

  • Programming

The Training-World Trap

A hypothesis can perform perfectly on the training set and still perform poorly on examples from the true data distribution. This is the central warning of overfitting: success on the observed sample does not necessarily mean success in the true world.

evaluates observed performanceevaluates true-world performanceTraining sampleobserved examplesHypothesisperformance can differTrue datadistributiontrue-world examples
What is contained in the finite training sample, what belongs to the underlying data distribution, and where can a hypothesis behave differently?

Overfitting is a mismatch between excellent training performance and poor true-distribution performance.

Two Ways to Measure Performance

Empirical error describes how a hypothesis performs on the observed training set. True error describes how that hypothesis performs on the true data distribution. The two measures can disagree because excellent performance on the training sample does not guarantee excellent performance in the true world.

training performancetrue-distribution performanceh_Sone hypothesisEmpirical error0True error1/2
How can a hypothesis classify every training example correctly while still making errors on examples drawn from the true data distribution?

Zero empirical error means that the hypothesis has no observed error on the training set. It does not mean that the hypothesis has zero true error.

A Perfect Training Score

The h_S Example

Interpret the case in which the empirical error of h_S is 0 while its true error is 1/2.

Read the empirical error: The value 0 means that h_S has no error on the training set.

Read the true error: The value 1/2 means that h_S has nonzero error on the true data distribution.

Compare the two results: The hypothesis is perfect on the observed training data but performs poorly on the true distribution.

Name the pattern: This mismatch is overfitting.

Training success alone is not enough to establish good true-distribution performance.

performanceperformanceh_Sone hypothesisTraining seterror 0True distributionerror 1/2
Which performance pattern indicates that h_S fits the observed training examples but performs poorly on unseen examples from the true distribution?

Why ERM Can Overfit

An empirical risk minimization algorithm selects a hypothesis because it minimizes empirical error. Its selection is therefore based on performance measured on the training data. If a hypothesis has the lowest observed training error but poor performance on the true data distribution, ERM may select that overfitting hypothesis.

comparerank observed errorchooseCandidatehypothesesmultiple choicesEmpirical errortraining dataLowest observed errorselection criterionSelected hypothesismay overfit
How does empirical risk minimization compare candidate hypotheses and select one with the lowest observed training error even when that hypothesis has higher true error?

ERM may choose an overfitting hypothesis not because its true error is lowest, but because its empirical error is lowest.

Following the Mismatch

To recognize overfitting, compare the hypothesis's performance on the training data with its performance on the true data distribution. The overfitting pattern appears when training performance is excellent while true-distribution performance is poor.

fit training sampleevaluate true distributionTraining errornot yet minimizedTraining error0True errornot specifiedTrue error1/2
What changes as a model fits the training sample more closely, and how can empirical error and true error differ?

Whenever a training result looks excellent, ask a second question: how does the hypothesis perform on the true data distribution? This comparison is necessary for detecting the mismatch that defines overfitting.

Mistakes in Error Interpretation

  • Treating zero empirical error as zero true error.

    The zero applies to the training set, while true error concerns the true data distribution.

    Fix: State precisely which error is zero and check the other measure separately.

  • Assuming the hypothesis selected by ERM must have the lowest true error.

    ERM minimizes empirical error, and the source example shows that empirical success can coexist with nonzero true error.

    Fix: Describe ERM's selection criterion as empirical error and avoid equating it automatically with true-distribution performance.

  • Ignoring the difference between the observed sample and the true data distribution.

    Overfitting is precisely the mismatch between these two performance settings.

    Fix: Name the evaluation setting whenever you report a performance result.

Check Your Interpretation

EASY

A hypothesis has empirical error 0 and true error 1/2. What does this tell you about its training performance, its true-distribution performance, and whether the pattern is overfitting?

Hints
  • Interpret empirical error using the training set.
  • Interpret true error using the true data distribution.
  • Compare the two results before naming the pattern.

What do you think happens?

If ERM selects a hypothesis with the lowest empirical error, can that selected hypothesis still be overfitting?

  • Yes
  • No
Reveal answer

Answer: Yes

ERM minimizes empirical error, but the selected hypothesis can still have poor performance on the true data distribution.

Key Takeaways

  1. Empirical error concerns performance on the observed training set.
  2. True error concerns performance on the true data distribution.
  3. Overfitting is the mismatch between excellent training performance and poor true-distribution performance.
  4. A hypothesis can have zero empirical error and nonzero true error, as in the example with true error 1/2.
  5. ERM may select an overfitting hypothesis because it minimizes empirical error.

Key Takeaways

  • Empirical error and true error measure performance in different settings.
  • Zero empirical error does not guarantee zero true error.
  • Overfitting occurs when training performance is excellent but true-distribution performance is poor.
  • ERM can select an overfitting hypothesis because its selection is based on empirical error.
  • The pattern L_S(h_S) = 0 and L_D(h_S) = 1/2 demonstrates why training success alone is not enough.