Inductive Bias
Finite hypothesis classes provide a controlled candidate set for ERM.
The Training-Data Trap
A learner can make no mistakes on the examples it has seen and still perform poorly on new examples. The reason is that success on a training sample is not automatically success over the underlying data distribution. Inductive bias addresses this problem by controlling which predictors a learner is allowed to consider.
Hypothesis Classes and ERM
A hypothesis class, written as H, is a set of predictors selected in advance. Each member h of H is a function that maps inputs from X to outputs in Y. Once H and the training sample S are available, empirical risk minimization, or ERM, chooses a member of H with the lowest error on S.
The class H is therefore more than a container for possible answers. It restricts the search performed by ERM. ERM remains responsible for comparing the permitted predictors, while H determines which predictors are permitted in the first place.
Choosing Among Permitted Predictors
Suppose H contains a finite collection of predictors, and the learner receives a training sample S.
Restrict the candidates: The learner may consider only the predictors that belong to H.
Measure sample performance: ERM evaluates how many mistakes each candidate makes on S.
Select the minimum: ERM chooses a member of H with the lowest empirical risk on S.
Ask the generalization question: The selected predictor must still be evaluated over the underlying data distribution, because the lowest training error does not by itself establish low true risk.
ERM selects the best candidate inside H according to the training sample, not necessarily the best predictor among every conceivable predictor.
Realizability and the Zero-Error Benchmark
The realizability assumption says that the target concept is represented by a hypothesis h* that belongs to H and has zero true risk. When a random training sample is drawn from the distribution and labeled by the target rule, h* has empirical risk zero on that sample with probability 1. Thus, under realizability, H contains at least one candidate that makes no errors on the observed training examples.
Because ERM chooses a hypothesis with minimum empirical risk, every ERM result hS also has empirical risk zero in this realizable setting. This is an important intermediate conclusion: it tells us what happens on the sample. It does not yet tell us that hS has zero true risk or even low true risk on the underlying distribution.
When Perfect Fit Fails
A learner that minimizes training error can overfit. Overfitting occurs when a predictor succeeds on the examples it has seen but performs poorly over the underlying data distribution. A highly flexible class may contain predictors that fit individual training examples rather than capturing a pattern that remains useful on new examples.
Class Size and Sample Size
For a finite hypothesis class, whether ERM generalizes successfully is tied to two conditions: the size of the class and the amount of training data. As the number of candidate hypotheses grows, the required amount of training data grows in relation to that class size. A larger class gives ERM more possible explanations to compare, so more data is needed to reliably identify a hypothesis with low true risk.
Finite does not necessarily mean extremely narrow. The source gives a class of predictors implementable by a C++ program using at most 10^9 bits of code as one example. It also notes that axis-aligned rectangles with unrestricted real-valued parameters form an infinite class, while representing those parameters with a 64-bit floating-point representation makes the resulting class finite.
Choosing the Bias Before Data
Choosing H before seeing the training data creates inductive bias. The choice expresses a prior structural belief about which predictors are appropriate for the task. After the class is fixed, ERM chooses among its members using the sample. The class therefore determines the space of explanations that ERM is allowed to consider.
Choose a hypothesis class using prior knowledge about the problem rather than treating it as an arbitrary technical container. Useful inductive bias should be motivated by the structure of the task.
Check Your Reasoning
A finite class H contains a hypothesis h* with zero true risk. ERM chooses hS with minimum empirical risk on a random realizable sample S. What can you conclude immediately, and what can you not conclude immediately?
Hints
- Use the fact that h* has empirical risk zero on S with probability 1.
- Remember that ERM chooses a minimum-empirical-risk member of H.
- Separate conclusions about S from conclusions about the underlying distribution.
What do you think happens?
If h* has zero true risk and belongs to H, what is the empirical risk of every ERM result hS under realizability?
Reveal answer
Answer: It is zero with probability 1.
The realizability assumption places h* in H with zero true risk, so h* has zero empirical risk on the random labeled sample with probability 1. ERM chooses a minimum-empirical-risk member of H, so its result also has empirical risk zero. This says nothing by itself about the result's true risk.
Treating zero empirical risk as proof of zero true risk.
The training sample is only part of the underlying data distribution, and a predictor can fit observed examples while performing poorly on new examples.
Fix:
State separately what is known about empirical risk on S and what must be established about true risk over the distribution.Assuming ERM searches through every possible predictor.
ERM selects only from the hypothesis class fixed for the learning problem.
Fix:
Say that ERM chooses a lowest-training-error member of H.Assuming a finite class automatically prevents overfitting.
The source states that the finite-set restriction alone is not enough to guarantee protection from overfitting.
Fix:
Consider both the quality of the restriction and the amount of training data.Ignoring class size when discussing data requirements.
The required amount of training data grows in relation to the size of the class.
Fix:
Connect larger candidate classes with a greater training-data requirement.
Key Takeaways
- A hypothesis class H is a set of predictors that restricts the choices available to ERM.
- ERM selects a member of H with minimum empirical risk on the training sample.
- Under realizability, H contains a zero-true-risk hypothesis, so ERM has a zero-empirical-risk candidate on the sample with probability 1.
- Zero empirical risk is not the same as low true risk; the difference is where overfitting can occur.
- The required training data grows with the size of the finite hypothesis class, and choosing H before seeing the data creates inductive bias.
Key Takeaways
- Inductive bias comes from choosing a hypothesis class before observing the training data.
- ERM searches only within that class and selects a predictor with minimum empirical risk.
- Realizability guarantees a zero-error candidate on the training sample with probability 1, but not automatically low true risk for the ERM result.
- Overfitting occurs when training performance fails to reflect performance over the underlying distribution.
- For finite classes, the amount of training data must grow in relation to the number of candidate hypotheses.