Concepts / Overfitting and Model Selection

Overfitting and Model Selection

SRM selects a hypothesis by minimizing empirical loss plus a penalty term.

  • Programming

Why Training Error Is Not Enough

Suppose several hypotheses are available for solving the same learning problem. One hypothesis may achieve a small loss on the training set, while another may have a slightly larger training loss but belong to a less complex hypothesis class. Structural Risk Minimization, or SRM, does not choose by empirical loss alone. It selects a hypothesis by considering the empirical loss together with a penalty associated with the hypothesis class containing that hypothesis.

The central idea is a balance: fit the observed training data, but also account for the hypothesis class used to obtain that fit.

contributescontributesEmpirical lossL_S(h)SRM selection valueloss plus penaltyClass penaltyassociated with H_n
How does SRM combine a hypothesis's training error with a complexity-related penalty when deciding which model is preferable?

The SRM Selection Rule

The source represents the overall hypothesis class as H = ⋃ₙ Hₙ. This means that the hypotheses considered by SRM are organized into a union of hypothesis classes. For a training set S and confidence parameter δ, the procedure searches through H and returns a hypothesis h that minimizes the displayed SRM objective. The objective contains the empirical loss L_S(h) and a penalty written in the source as ε_n(h)(m, w(n(h)) · δ).

The notation matters because the selection is not described as choosing the hypothesis with the smallest training loss in isolation. The selected hypothesis minimizes the combined value specified by the SRM rule. A hypothesis with a particularly small empirical loss can therefore be compared with another hypothesis whose combined loss-and-penalty value is better.

search withcontainsevaluateminimizeS and δprocedure inputsH = union of H_noverall hypothesis classHypothesis hcandidate from H_nLoss plus penaltySRM objectiveSelected hminimum displayed sum
How does SRM compare candidate hypotheses across the union of hypothesis classes and select one with the lowest penalized empirical loss?

Uniform Convergence Within Classes

Each H_n in the SRM structure has uniform convergence. In this setting, uniform convergence connects empirical loss and true risk for every hypothesis in a class at the same time, rather than describing a guarantee for only one hypothesis considered in isolation. This property is part of the structure in which SRM operates.

This gives the class structure a specific role. SRM does not merely list unrelated candidate hypotheses; it organizes them into classes H_n for which uniform convergence holds. The selection rule then uses the empirical loss of a candidate together with the penalty associated with the class in which that candidate lies.

containshavehaveconnected toconnected toH_nhypothesis classHypotheses in H_nmany candidatesEmpirical losstraining-set behaviorUniform convergenceclass-wide connectionTrue riskgeneralization behavior
How does uniform convergence connect empirical loss and true risk for every hypothesis in a class at the same time?

Following Complexity Through the Choice

Consider a generated comparison involving two candidate classes. A hypothesis from one class has a very small empirical loss. A hypothesis from another class has a different empirical loss and a different class-associated penalty. SRM does not decide from the first loss value alone. It evaluates the displayed combined objective for both candidates and selects the hypothesis with the smaller combined value.

Comparing Two Candidate Hypotheses

Two hypotheses are being considered under SRM. Hypothesis h1 has a smaller empirical loss, while hypothesis h2 has a different penalty associated with its class. What information must be compared before selecting a hypothesis?

Identify the training contribution: Record the empirical loss L_S(h) for each candidate. This is the part of SRM that measures behavior on the training set.

Identify the class contribution: Record the penalty associated with the hypothesis class containing each candidate. SRM includes this contribution in addition to empirical loss.

Form the SRM comparison: Compare the combined empirical-loss-and-penalty value for h1 and h2 rather than comparing empirical losses by themselves.

Apply the selection rule: Choose the hypothesis that minimizes the displayed SRM sum over the overall hypothesis class.

The correct SRM choice cannot be determined from training loss alone. It is the candidate with the lower combined value under the SRM rule.

evaluateevaluateCandidate hclass-associated penaltyEmpirical loss pluspenaltySRM comparison valueCandidate hdifferent class-associatedpenaltyEmpirical loss pluspenaltyupdated SRM comparisonvalue
What changes in the SRM comparison when a candidate comes from a more complex hypothesis class?

How SRM Addresses Overfitting

Overfitting is the problem SRM is intended to address through its two-part objective. Empirical loss minimization is included, but it is not the whole principle. Because SRM also includes a penalty associated with the hypothesis class, choosing a hypothesis requires considering more than how closely it fits the training set.

The safe interpretation is a selection principle, not a guaranteed outcome for every possible model collection. SRM searches the overall class H and chooses the hypothesis minimizing the specified sum. The penalty discourages treating the smallest training loss as the only criterion, thereby providing a mechanism intended to reduce the risk of selecting an overfitted hypothesis.

does not include penaltyminimize togetheraddressed by addingEmpirical loss onlyone contribution consideredOverfitting riskintended problem to addressEmpirical loss pluspenaltySRM objectiveSRM-selectedhypothesisminimum combined value
How does the complexity penalty discourage choosing an overly flexible hypothesis that fits training data but may overfit?

Mistakes in Reading SRM

  • Treating SRM as ordinary empirical loss minimization

    Empirical loss minimization is only one part of SRM.

    Fix: Compare the combined empirical loss and penalty specified by the SRM rule.

  • Ignoring the union of hypothesis classes

    The overall hypothesis class is represented as H = ⋃ₙ Hₙ.

    Fix: Interpret the procedure as searching through the overall class H, which is formed from the SRM classes.

  • Confusing uniform convergence with the penalty itself

    Uniform convergence is a property of each H_n in the SRM structure; the penalty is a separate contribution in the selection objective.

    Fix: Keep the class property and the selection penalty conceptually distinct.

  • Assuming the source specifies an exact penalty curve

    The source does not provide a numerical curve or a rule for every possible model.

    Fix: State only that SRM includes a penalty associated with the hypothesis class.

Check Your Interpretation

MEDIUM

A learning procedure searches H = ⋃ₙ Hₙ. It receives a training set S and confidence parameter δ. It evaluates each candidate using empirical loss together with the penalty associated with the candidate's class. In your own words, explain why this procedure is SRM rather than empirical loss minimization alone.

Hints
  • Name both contributions in the objective.
  • Explain what the union H = ⋃ₙ Hₙ represents.
  • Mention how the penalty is related to overfitting.

What do you think happens?

If two hypotheses have different empirical losses, is the one with the smaller empirical loss automatically the SRM choice?

  • Yes, always
  • No, the combined value is compared
  • Only if both hypotheses are identical
Reveal answer

Answer: No. SRM selects by minimizing the combined empirical loss and class-associated penalty.

Empirical loss is part of the SRM objective, but it is not the whole principle. The penalty must also be included in the comparison.

The Selection Principle

  1. SRM represents the overall hypothesis class as H = ⋃ₙ Hₙ.
  2. Each H_n in the SRM structure has uniform convergence, connecting empirical loss and true risk for every hypothesis in that class at the same time.
  3. SRM searches the overall hypothesis class using the training set S and confidence parameter δ.
  4. The selected hypothesis minimizes empirical loss together with a penalty associated with its hypothesis class.
  5. This penalty-based comparison is intended to reduce the risk of overfitting by preventing training loss from being the only selection criterion.

Key Takeaways

  • Structural Risk Minimization selects a hypothesis by minimizing empirical loss plus a class-associated penalty.
  • The overall hypothesis class is organized as H = ⋃ₙ Hₙ, and each H_n has uniform convergence.
  • Uniform convergence connects empirical loss and true risk across all hypotheses in a class at the same time.
  • The SRM penalty makes model selection more than a search for the smallest training loss.
  • SRM is intended to reduce overfitting by including class complexity in the selection decision.