Concepts / True Risk and Empirical Risk

True Risk and Empirical Risk

Overfitting occurs when a hypothesis fits the training data too well but performs poorly on the true distribution.

  • Programming

The Hidden Question Behind Zero Error

A hypothesis with no mistakes on its training set can look successful. But the training set is only the data used to choose the hypothesis. The more important question is how that hypothesis behaves on the true distribution that produces instances. Overfitting occurs when a hypothesis fits the training data too well but performs poorly on the true distribution.

evaluated onevaluated onTraining samplelow observed errorOverfittinghypothesisfits sample too wellTrue distributionhigh broader error
How can a hypothesis have low error on the observed training sample but high error on the true data distribution?

Two Ways to Measure Performance

Empirical risk describes performance computed from the finite training sample: it asks how well the hypothesis handles the instances already observed. True risk describes performance on the true distribution: it asks how well the hypothesis handles the broader world represented by the distribution. These are different questions, even when they are applied to the same hypothesis.

evaluated ondeterminesevaluated ondeterminesHypothesisTraining sampleobserved instancesEmpirical risksample performanceTrue distributioninstance-producing processTrue riskdistribution performance
How are empirical risk, computed from a finite sample, and true risk, averaged over the full data distribution, connected?

The Papaya Prediction Problem

Consider predicting whether a papaya tastes good using two features: softness and color. In the source's geometric construction, instances are spread uniformly through a gray square. The true labeling rule assigns label 1 to points inside an inner blue square and label 0 to points outside it. The gray square has area 2, while the blue square has area 1.

A Perfect Training Result with Poor Distribution Performance

What does the papaya construction show about a hypothesis that makes no mistakes on the training sample?

Observe the sample result: The source describes a predictor with zero training error. On the observed training examples, it appears completely successful.

Evaluate beyond the sample: The same predictor is evaluated against the true distribution that produces papaya instances, rather than only against the observed examples.

Compare the two risks: The source states that this same predictor has true error equal to one half, despite having zero training error.

Zero empirical error does not guarantee low true error. The predictor is an example of overfitting because its training performance is perfect while its performance on the true distribution is poor.

evaluatesevaluatesTraining exampleszero errorSame predictorTrue distributiontrue error one half
What changes when a hypothesis moves from evaluating the training examples to evaluating new examples drawn from the true distribution?

Why ERM Can Overfit

The empirical risk minimization rule relies on training-set performance. It selects a hypothesis according to how well that hypothesis performs on the observed training data. Therefore, if a memorizing hypothesis has the lowest observed training error, ERM can select it even when another hypothesis captures the underlying labeling pattern and performs better on the true distribution.

competes on sample errorcompetes on sample errorselectsMemorizinghypothesislowest sample errorERM ruleuses training performanceSelected hypothesismemorizerPattern hypothesisbetter true-distributionperformance
How can empirical risk minimization choose a memorizing hypothesis because it has the lowest observed training error, even when another hypothesis performs better on unseen data?

Memorizing or Learning the Pattern

A hypothesis that memorizes the observed sample can match the labels of the examples it has already seen without capturing the rule that assigns labels more broadly. A hypothesis that captures the underlying labeling pattern aims to reflect how labels are assigned throughout the true distribution. The distinction matters because training performance examines the observed examples, while true-distribution performance examines behavior beyond those examples.

fitsdescribes labels forObserved examplestraining sampleSample memorizationmatches observed labelsNew instancestrue distributionUnderlying labelingpatternguides broader predictions
What is the difference between a hypothesis that memorizes each observed example and one that learns the rule generating labels for new examples?

Whenever a hypothesis has no training mistakes, delay the conclusion that it is successful. Ask a second question: how does it perform on the true distribution that produces instances?

Mistakes About Training Success

  • Treating zero empirical error as proof of low true error.

    The source gives a predictor with zero training error and true error equal to one half.

    Fix: Separate performance on the observed sample from performance on the true distribution.

  • Assuming the ERM-selected hypothesis must capture the underlying labeling pattern.

    The ERM rule relies on training-set performance, so it can select an overfitting hypothesis.

    Fix: Recognize that the lowest observed error can belong to a hypothesis that memorizes the sample.

  • Using the words training performance and true-distribution performance as if they described the same evaluation.

    The training set is only the data used to choose the hypothesis; the true distribution represents the broader source of instances.

    Fix: Name which set or distribution is being used whenever you discuss error.

Check Your Understanding

MEDIUM

A hypothesis has zero error on the observed training sample. Explain why that fact alone is insufficient to conclude that the hypothesis generalizes well. Then describe how ERM could still select this hypothesis.

Hints
  • Identify what empirical risk evaluates.
  • Contrast the training sample with the true distribution.
  • Consider a hypothesis that memorizes the observed examples.

What do you think happens?

A predictor has zero training error. What can you conclude about its true error from that fact alone?

  • It must also have zero true error.
  • It must have low true error.
  • Nothing definitive about low true error follows from the training result alone.
Reveal answer

Answer: Nothing definitive about low true error follows from the training result alone.

The source explicitly states that zero empirical error does not guarantee low true error and gives a predictor whose true error is one half.

Key Takeaways

  1. Empirical risk measures how a hypothesis performs on the observed training sample.
  2. True risk measures how the hypothesis performs on the true distribution that produces instances.
  3. Overfitting occurs when training performance is unusually favorable while performance on the true distribution is poor.
  4. Zero empirical error does not guarantee low true error.
  5. Because ERM relies on training-set performance, it can select a hypothesis that memorizes the sample instead of capturing the underlying labeling pattern.

Key Takeaways

  • Empirical risk concerns the observed training sample, whereas true risk concerns the true distribution.
  • A hypothesis can achieve zero training error and still perform poorly beyond the training data.
  • ERM can choose an overfitting hypothesis because it selects according to training-set performance.
  • Memorizing observed examples is different from capturing the underlying labeling pattern.