Concepts / Overfitting and Generalization

Overfitting and Generalization

Finite hypothesis classes provide a controlled candidate set for ERM.

  • Programming

A Perfect Training Score

Imagine two hypotheses that both classify every example in a training sample correctly. One may have learned a pattern that continues to work on new data. The other may have matched accidental details of the sample and perform poorly beyond it. This is the central difficulty of overfitting: success on the observed sample does not by itself establish success under the data distribution.

evaluated onevaluated onTraining sampleobserved examplesZero empirical riskno sample errorsUnseen datadata distributionHigh true riskmany distribution errors
How can a hypothesis make no mistakes on the training sample yet still make many mistakes on unseen data?

The ERM Selection Process

Suppose a learner may choose only from a fixed collection of predictors, called a hypothesis class. In this article, the class contains a finite number of candidates. Empirical risk minimization, or ERM, evaluates how each candidate performs on the training sample and selects a hypothesis with minimum empirical risk. ERM therefore turns the training sample into a choice among the available candidates.

provides candidatesprovides evidenceproduces risksselects minimumFinite hypothesisclasscandidate predictorsTraining sampleobserved dataEvaluate candidatessample performanceCompare empiricalrisksfind minimumERM resultminimum empirical risk
How does ERM compare candidate hypotheses and select the one with the lowest empirical risk?

Selecting among three candidates

An ERM learner compares three hypotheses on the same training sample. Hypothesis A makes errors on the sample, while hypotheses B and C make no errors.

Evaluate: ERM examines the performance of A, B, and C on the training sample.

Compare: A has higher empirical risk than B and C. B and C are tied for the minimum empirical risk.

Select: ERM chooses a hypothesis with minimum empirical risk, so it selects B or C in this example.

The ERM result has zero empirical risk on the sample, but this result alone does not establish low true risk.

Realizability and the Sample

The realizability assumption places a target hypothesis h* inside the hypothesis class H and gives h* zero true risk. In other words, the class contains a hypothesis that agrees with the target labeling rule in the setting under discussion.

Because h* has zero true risk, a random sample drawn from the data distribution and labeled by the target rule has empirical risk zero for h* with probability 1. Thus, under realizability, the class contains at least one candidate that labels every training example correctly. ERM can use this candidate as a benchmark.

belongs tohasimplies correct labels ongives h*Target ruleh*Hypothesis classHZero true riskRandom labeled sampleZero empirical risk
What does it mean for the target labeling rule to belong to the hypothesis class, and why does that imply that some hypothesis labels every training example correctly?

From Sample Fit to Generalization

Under realizability, h* has zero empirical risk on the random training sample with probability 1. Since ERM chooses a hypothesis with minimum empirical risk, every ERM result h_S also has empirical risk zero in this setting. This is an important intermediate conclusion: ERM can fit the observed sample perfectly.

The conclusion is still limited to the sample. The fact that h_S has zero empirical risk does not by itself prove that h_S has low true risk. A selected hypothesis can agree with all observed examples while reflecting accidental patterns that do not continue beyond the sample. Generalization asks whether the training-based choice also performs well under the data distribution.

Two hypotheses with the same training score

Consider two candidates selected after examining a training sample. Both classify every training example correctly. Candidate A follows a pattern that also agrees with the target rule on unseen data. Candidate B matches the training examples but uses accidental details that do not continue on unseen data.

On the training sample: Both candidates have zero empirical risk because neither makes an error on the observed examples.

Beyond the sample: The candidates can differ in true risk because their behavior under the data distribution need not be the same.

Interpret the result: Zero empirical risk identifies perfect sample performance, not automatically good generalization.

A learner must reason about the hypothesis class and the amount of training data, not only the training score.

fitsfitsmay havemay haveTraining samplesame observed examplesHypothesis Azero empirical riskLow true riskunseen dataHypothesis Bzero empirical riskHigh true riskunseen data
How can two hypotheses with identical training performance differ when evaluated under the data distribution?

Why Class Size Matters

A finite hypothesis class gives ERM a controlled candidate set. Finiteness does not necessarily mean that the modeling idea is extremely narrow. For example, the source material describes a class of predictors implementable by a C++ program using at most 10^9 bits of code. It also describes axis-aligned rectangles: the class is infinite when real-valued parameters have no finite representation limit, but discretizing those parameters with a 64-bit floating-point representation makes the resulting class finite.

Restricting the learner to a finite collection limits the number of candidate explanations that ERM can compare and select from. This control can limit overfitting because the learner is not choosing from an unbounded set of possible predictors. The required amount of training data is tied to the size of the class: as the class becomes larger, the required training-sample size grows in relation to that class size.

is tied tois tied toSmaller classfewer candidatesRequired datasmaller amountLarger classmore candidatesRequired datalarger amount
Why does ERM need more training examples when it chooses from a larger hypothesis class?
providesallows moreFinite candidatesetcontrolled choicesControlled ERM searchlimits possible fitsUnbounded candidatechoicesmore possible fitsAccidental patternsmore ways to fit
How does restricting the learner to a finite set of candidate hypotheses reduce the number of ways it can fit accidental patterns in the training data?

Common Reasoning Errors

  • Treating zero empirical risk as zero true risk.

    Empirical risk concerns the training sample, while true risk concerns performance under the data distribution.

    Fix: State the conclusion precisely: the hypothesis has zero empirical risk on the sample. Generalization still requires reasoning about true risk.

  • Assuming ERM selects the universally correct hypothesis.

    ERM selects a hypothesis with minimum empirical risk. Multiple candidates can have the same sample performance, and the sample result alone does not identify the target hypothesis.

    Fix: Describe ERM as a training-sample selection procedure, then separately ask whether its result generalizes.

  • Forgetting the realizability assumption when claiming zero empirical risk for ERM.

    The zero-empirical-risk conclusion follows here because realizability places h* in the class and gives it zero true risk.

    Fix: Identify the assumption first: under realizability, h* has zero empirical risk on the random labeled sample with probability 1.

  • Interpreting a finite class as necessarily very small.

    A finite class can still contain many candidates. The source gives examples involving bounded-length C++ programs and discretized axis-aligned rectangles.

    Fix: Use finite to mean that the candidate collection has a finite number of members, not that the collection is tiny.

  • Ignoring class size when discussing the amount of data needed.

    The required amount of training data grows in relation to the size of the hypothesis class.

    Fix: Always connect the data requirement to the number of available hypotheses.

Check Your Understanding

MEDIUM

A finite hypothesis class contains a target hypothesis h* with zero true risk. A random sample is labeled by the target rule, and ERM selects a hypothesis with minimum empirical risk. What can you conclude about the ERM result on the training sample? What can you not conclude from that fact alone?

Hints
  • Use the fact that h* belongs to the class and has zero true risk.
  • Remember that ERM chooses a minimum empirical-risk hypothesis.
  • Separate performance on the observed sample from performance under the data distribution.

Checking the conclusion

Explain the consequence of realizability for ERM and distinguish it from a generalization claim.

Locate h*: Realizability places h* inside H and gives h* zero true risk.

Transfer to the sample: A random sample labeled by the target rule has empirical risk zero for h* with probability 1.

Apply ERM: Because ERM chooses a minimum empirical-risk hypothesis, its result also has empirical risk zero in this setting.

Limit the conclusion: This establishes perfect performance on the observed sample, not automatically low true risk under the data distribution.

Under realizability, ERM has access to a zero-error candidate on the sample and therefore selects a zero-empirical-risk result, but generalization remains the separate question.

Key Takeaways

  1. ERM examines a finite collection of candidate hypotheses and selects one with minimum empirical risk on the training sample.
  2. Under realizability, the class contains h* with zero true risk, so h* has zero empirical risk on a random labeled sample with probability 1.
  3. Consequently, every ERM result has zero empirical risk in this realizable finite-class setting.
  4. Zero empirical risk concerns the observed sample and is not the same as low true risk under the data distribution.
  5. The required amount of training data grows in relation to the size of the hypothesis class, which is why controlling the candidate set can limit overfitting.

Key Takeaways

  • ERM selects a minimum-empirical-risk hypothesis from a finite candidate class.
  • Realizability places a zero-true-risk target hypothesis inside that class.
  • Under realizability, ERM achieves zero empirical risk on the random training sample with probability 1.
  • Zero empirical risk does not by itself guarantee low true risk or good generalization.
  • Larger finite hypothesis classes require a larger amount of training data, so restricting the candidate set can help control overfitting.