Concepts / Generalization and Learnability

Generalization and Learnability

ϵ-representativeness describes a condition on a training set S relative to Z, H, ℓ, and D.

  • Programming

Why the Sample Matters

A learning rule receives a training set and uses it to choose a hypothesis. The rule does not examine the entire learning problem directly; it relies on the information contained in the training set. If the training set reflects the learning problem well, the rule has a better basis for choosing a useful hypothesis. If it does not, even a sensible learning rule may be working with an unreliable picture.

ϵ-representativeness is a formal condition on the training set. It describes how closely the losses measured on the training set correspond to losses under the distribution, across the hypothesis class.

The Four Ingredients

A sample is not simply representative by itself. Its representativeness is always understood relative to four parts of a learning setting: a domain Z, a hypothesis class H, a loss function ℓ, and a distribution D. Changing any of these ingredients changes the learning setting in which representativeness is being judged.

defines settingdefines comparisonsmeasures errordefines population viewsatisfies conditionDomain ZTraining set Srepresentativeness isevaluated hereϵ-representativerelative to Z, H, ℓ, and DHypothesis class HLoss function ℓDistribution D
How do the domain, hypotheses, loss, distribution, and training set fit together?

A training set S is ϵ-representative, relative to Z, H, ℓ, and D, when the losses measured from S and the corresponding losses under D are sufficiently close for the hypotheses in H, according to the definition of ϵ-representativeness.

Uniform Loss Comparison

The important word in this idea is uniform. Representativeness is not a claim about only the hypothesis eventually selected by the learning rule. It compares the training-set loss with the distribution-based loss for every hypothesis in H. The sample is useful when it gives a consistently reliable picture across the whole hypothesis class.

measuresmeasuresmeasuresmeasurescompare corresponding lossescompare corresponding lossesTraining set Sempirical lossHypothesis h1loss on SDistribution Dexpected lossHypothesis h1loss under DHypothesis h2loss on SHypothesis h2loss under D
For every hypothesis in H, how closely does the training-set loss match the loss under D?

How ERM Uses the Guarantee

Empirical risk minimization, or ERM, chooses a hypothesis with the lowest loss measured on the training set. That choice is based on empirical information. The reason (ϵ/2)-representativeness matters is that it makes the empirical comparison sufficiently close to the corresponding comparison under D. When the sample is (ϵ/2)-representative, the ERM learning rule is guaranteed to return a good hypothesis.

supports uniform closenessmakes empirical choice informativeguaranteesSample S(ϵ/2)-representativeLoss comparisontraining and distributionERM choicelowest empirical lossReturned hypothesisgood hypothesis
How does the (ϵ/2) gap between empirical and distribution-based loss combine with ERM's lowest-loss choice?

Why the empirical winner can be trusted

Suppose a hypothesis class contains several candidates, and ERM selects the candidate with the lowest loss on training set S. What role does (ϵ/2)-representativeness play?

Start with the sample: Treat S as a sample whose representativeness is being evaluated relative to the domain, hypothesis class, loss function, and distribution.

Compare all candidates: Because representativeness is uniform, the training-set loss and the corresponding loss under D are compared across the hypotheses in H, not only for the candidate ERM eventually selects.

Apply ERM: ERM chooses a hypothesis with the lowest empirical loss. The uniform comparison makes this training-set choice informative about the corresponding losses under D.

Use the guarantee: When S is (ϵ/2)-representative, the stated ERM guarantee applies: the learning rule returns a good hypothesis.

The sample's (ϵ/2)-representativeness supports the quality guarantee for the hypothesis selected by ERM.

The guarantee comes from the combination of two facts: representativeness makes empirical losses reliable across H, and ERM selects according to those empirical losses.

Sample Property and Hypothesis Quality

These are related but different statements. Saying that S is ϵ-representative describes the training set and how accurately it reflects the learning setting. Saying that the hypothesis returned by ERM is good describes the result of applying a learning rule to that sample. The first statement concerns the input to the learning rule; the second concerns its output.

QuestionObject being describedMeaning
Is S ϵ-representative?Training set SThe sample accurately compares training-set and distribution-based losses relative to the learning setting.
Is the ERM output good?Hypothesis returned by ERMThe selected hypothesis has the quality guaranteed when the required representativeness condition holds.

Representativeness and returned-hypothesis quality are connected, but they are not the same property.

When reading or writing a proof about ERM, name the object being discussed. Use sample language for representativeness and hypothesis language for the quality of the ERM output.

Mistakes to Avoid

  • Treating representativeness as a property of a sample in isolation.

    Representativeness is defined relative to those four ingredients of the learning setting.

    Fix: State the domain, hypothesis class, loss function, and distribution whenever the representativeness condition matters.

  • Checking only the hypothesis selected by ERM.

    The condition is uniform over the hypotheses in H.

    Fix: Understand ϵ-representativeness as a comparison that applies across the hypothesis class.

  • Confusing low empirical loss with the definition of representativeness.

    Low loss for a selected hypothesis describes a result on S; representativeness describes how training-set and distribution-based losses relate across H.

    Fix: Keep the sample property and the selected hypothesis property separate.

  • Forgetting the role of (ϵ/2)-representativeness in the ERM result.

    The guarantee stated for ERM depends on S being (ϵ/2)-representative.

    Fix: Connect the ERM guarantee explicitly to the (ϵ/2)-representativeness assumption.

Check Your Understanding

MEDIUM

In your own words, explain why a sample cannot be called ϵ-representative without reference to Z, H, ℓ, and D. Then explain the difference between saying that S is (ϵ/2)-representative and saying that ERM returns a good hypothesis.

Hints
  • Start by naming the four ingredients of the learning setting.
  • Describe representativeness as a comparison involving losses for hypotheses in H.
  • Treat the sample condition as the premise and the quality of the ERM output as the resulting guarantee.

What do you think happens?

If a sample is known to be (ϵ/2)-representative and ERM chooses the hypothesis with the lowest empirical loss, what does the stated guarantee say about the returned hypothesis?

Reveal answer

Answer: The ERM learning rule is guaranteed to return a good hypothesis.

The representativeness condition makes the empirical comparisons informative across the hypothesis class, and the source states that (ϵ/2)-representativeness guarantees a good ERM output.

Key Takeaways

  1. ϵ-representativeness is a condition on a training set, not a label that makes sense without context.
  2. The relevant context consists of the domain Z, hypothesis class H, loss function ℓ, and distribution D.
  3. Representativeness compares training-set loss with distribution-based loss uniformly across the hypotheses in H.
  4. When S is (ϵ/2)-representative, ERM is guaranteed to return a good hypothesis.
  5. The representativeness of S and the quality of the hypothesis returned by ERM are distinct statements.

Key Takeaways

  • ϵ-representativeness describes how well a training set reflects the learning problem.
  • The definition is relative to Z, H, ℓ, and D.
  • The comparison applies uniformly to every hypothesis in H.
  • (ϵ/2)-representativeness enables the ERM guarantee because ERM's empirical choice is supported by a reliable loss comparison.
  • A property of the sample should not be confused with the quality of the hypothesis selected from that sample.