Generalization and Learnability
ϵ-representativeness describes a condition on a training set S relative to Z, H, ℓ, and D.
Why the Sample Matters
A learning rule receives a training set and uses it to choose a hypothesis. The rule does not examine the entire learning problem directly; it relies on the information contained in the training set. If the training set reflects the learning problem well, the rule has a better basis for choosing a useful hypothesis. If it does not, even a sensible learning rule may be working with an unreliable picture.
ϵ-representativeness is a formal condition on the training set. It describes how closely the losses measured on the training set correspond to losses under the distribution, across the hypothesis class.
The Four Ingredients
A sample is not simply representative by itself. Its representativeness is always understood relative to four parts of a learning setting: a domain Z, a hypothesis class H, a loss function ℓ, and a distribution D. Changing any of these ingredients changes the learning setting in which representativeness is being judged.
A training set S is ϵ-representative, relative to Z, H, ℓ, and D, when the losses measured from S and the corresponding losses under D are sufficiently close for the hypotheses in H, according to the definition of ϵ-representativeness.
Uniform Loss Comparison
The important word in this idea is uniform. Representativeness is not a claim about only the hypothesis eventually selected by the learning rule. It compares the training-set loss with the distribution-based loss for every hypothesis in H. The sample is useful when it gives a consistently reliable picture across the whole hypothesis class.
How ERM Uses the Guarantee
Empirical risk minimization, or ERM, chooses a hypothesis with the lowest loss measured on the training set. That choice is based on empirical information. The reason (ϵ/2)-representativeness matters is that it makes the empirical comparison sufficiently close to the corresponding comparison under D. When the sample is (ϵ/2)-representative, the ERM learning rule is guaranteed to return a good hypothesis.
Why the empirical winner can be trusted
Suppose a hypothesis class contains several candidates, and ERM selects the candidate with the lowest loss on training set S. What role does (ϵ/2)-representativeness play?
Start with the sample: Treat S as a sample whose representativeness is being evaluated relative to the domain, hypothesis class, loss function, and distribution.
Compare all candidates: Because representativeness is uniform, the training-set loss and the corresponding loss under D are compared across the hypotheses in H, not only for the candidate ERM eventually selects.
Apply ERM: ERM chooses a hypothesis with the lowest empirical loss. The uniform comparison makes this training-set choice informative about the corresponding losses under D.
Use the guarantee: When S is (ϵ/2)-representative, the stated ERM guarantee applies: the learning rule returns a good hypothesis.
The sample's (ϵ/2)-representativeness supports the quality guarantee for the hypothesis selected by ERM.
The guarantee comes from the combination of two facts: representativeness makes empirical losses reliable across H, and ERM selects according to those empirical losses.
Sample Property and Hypothesis Quality
These are related but different statements. Saying that S is ϵ-representative describes the training set and how accurately it reflects the learning setting. Saying that the hypothesis returned by ERM is good describes the result of applying a learning rule to that sample. The first statement concerns the input to the learning rule; the second concerns its output.
| Question | Object being described | Meaning |
|---|---|---|
| Is S ϵ-representative? | Training set S | The sample accurately compares training-set and distribution-based losses relative to the learning setting. |
| Is the ERM output good? | Hypothesis returned by ERM | The selected hypothesis has the quality guaranteed when the required representativeness condition holds. |
Representativeness and returned-hypothesis quality are connected, but they are not the same property.
When reading or writing a proof about ERM, name the object being discussed. Use sample language for representativeness and hypothesis language for the quality of the ERM output.
Mistakes to Avoid
Treating representativeness as a property of a sample in isolation.
Representativeness is defined relative to those four ingredients of the learning setting.
Fix:
State the domain, hypothesis class, loss function, and distribution whenever the representativeness condition matters.Checking only the hypothesis selected by ERM.
The condition is uniform over the hypotheses in H.
Fix:
Understand ϵ-representativeness as a comparison that applies across the hypothesis class.Confusing low empirical loss with the definition of representativeness.
Low loss for a selected hypothesis describes a result on S; representativeness describes how training-set and distribution-based losses relate across H.
Fix:
Keep the sample property and the selected hypothesis property separate.Forgetting the role of (ϵ/2)-representativeness in the ERM result.
The guarantee stated for ERM depends on S being (ϵ/2)-representative.
Fix:
Connect the ERM guarantee explicitly to the (ϵ/2)-representativeness assumption.
Check Your Understanding
In your own words, explain why a sample cannot be called ϵ-representative without reference to Z, H, ℓ, and D. Then explain the difference between saying that S is (ϵ/2)-representative and saying that ERM returns a good hypothesis.
Hints
- Start by naming the four ingredients of the learning setting.
- Describe representativeness as a comparison involving losses for hypotheses in H.
- Treat the sample condition as the premise and the quality of the ERM output as the resulting guarantee.
What do you think happens?
If a sample is known to be (ϵ/2)-representative and ERM chooses the hypothesis with the lowest empirical loss, what does the stated guarantee say about the returned hypothesis?
Reveal answer
Answer: The ERM learning rule is guaranteed to return a good hypothesis.
The representativeness condition makes the empirical comparisons informative across the hypothesis class, and the source states that (ϵ/2)-representativeness guarantees a good ERM output.
Key Takeaways
- ϵ-representativeness is a condition on a training set, not a label that makes sense without context.
- The relevant context consists of the domain Z, hypothesis class H, loss function ℓ, and distribution D.
- Representativeness compares training-set loss with distribution-based loss uniformly across the hypotheses in H.
- When S is (ϵ/2)-representative, ERM is guaranteed to return a good hypothesis.
- The representativeness of S and the quality of the hypothesis returned by ERM are distinct statements.
Key Takeaways
- ϵ-representativeness describes how well a training set reflects the learning problem.
- The definition is relative to Z, H, ℓ, and D.
- The comparison applies uniformly to every hypothesis in H.
- (ϵ/2)-representativeness enables the ERM guarantee because ERM's empirical choice is supported by a reliable loss comparison.
- A property of the sample should not be confused with the quality of the hypothesis selected from that sample.