Concepts / Generalization from Training Data

Generalization from Training Data

ERM can overfit because success on a training sample does not ensure success over the underlying data distribution.

  • Programming

A Perfect Training Score

A learner can inspect a training sample and choose a predictor that makes the fewest mistakes on those examples. This is a natural strategy, but it has a serious limitation: success on the examples already seen does not guarantee success on new examples from the underlying data distribution. When a predictor fits the training sample well but performs poorly over that broader distribution, the failure is called overfitting.

does not guaranteeTraining samplelow errorData distributionpoor performance
How can a hypothesis fit every training example yet perform poorly on new examples drawn from the underlying distribution?

Tracing the Learning Input

Training data is the learner’s input. It is a finite sequence of labeled examples written as S = ((x₁, y₁), ..., (xₘ, yₘ)). Each pair belongs to X × Y: xᵢ is an input, or domain point, from X, and yᵢ is its label from Y. The individual pairs are training examples, while the complete sequence S is also called the training set.

containscontainspairspairspairspairsTraining data Sfinite sequenceTraining example(x₁, y₁)Input x₁from XInput x₂from XLabel y₁from YTraining example(x₂, y₂)Label y₂from Y
How is a training-data sequence organized into examples, and how does each example contain an input from X paired with a label from Y?

This labeling step is essential. A collection containing only inputs from X gives the learner domain points, but it does not provide the associated labels from Y. The collection becomes labeled training data when each input is paired with its label.

containscontainsLabeled data(xᵢ, yᵢ)InputxᵢLabelyᵢUnlabeled dataxᵢ only
What information is present in labeled training data that is missing from a collection containing only inputs?

Restricting ERM with H

Empirical risk minimization, or ERM, chooses a predictor with the lowest possible error over the training data. If ERM may consider every available predictor, it may find one that performs extremely well on the observed sample without performing well over the underlying data distribution.

A hypothesis class is a set of predictors chosen in advance. It is written as H. Each member h of H is a function that maps inputs from X to outputs in Y. After H has been fixed and the training sample S is available, ERM operates inside H: it chooses a member of H with the lowest error on S. The class therefore restricts which candidate predictors ERM is allowed to consider.

restricts tocontainscontainscandidatecandidateAvailablepredictorsHypothesis class HPredictor h₁member of HERM choicelowest error on SPredictor h₂member of H
How does choosing a hypothesis class change which candidate predictors ERM is allowed to consider?

Choosing Within a Hypothesis Class

Suppose a learner has a training sample S and a hypothesis class H containing several permitted predictors.

Set the permitted choices: The learner fixes H, so ERM is not allowed to search through every available predictor.

Observe the sample: The learner receives S, a finite sequence of labeled examples.

Compare permitted predictors: ERM evaluates the predictors that belong to H by their errors on S.

Select a member: ERM chooses a member of H with the lowest error on S.

The hypothesis class determines which predictors count as permitted choices; ERM chooses among those permitted predictors.

Inductive Bias Before Observation

Choosing H before seeing the training data creates inductive bias. The choice expresses a prior structural belief about which predictors are appropriate for the task. ERM remains responsible for selecting among the permitted predictors, while H determines what counts as a permitted choice.

thenwith H fixedERMChoose Hbefore SObserve Slabeled sampleRestrict choicesmembers of HSelect hlowest error on S
What happens in the learning process when the hypothesis class is chosen before the training sample is observed?

The order matters. If the class is selected before the sample is observed, its restriction is not merely a reaction to the particular examples in S. It is a prior choice about the kinds of predictors the learner will permit. Useful inductive bias should be motivated by prior knowledge about the problem.

A Concrete Labeling Example

Consider a collection of papayas described by color and softness. If the associated tastiness result is recorded for each described papaya, each item is a labeled pair: the description supplies the input and tastiness supplies the label. The resulting collection can serve as training data. If only color and softness are collected, the inputs are present but the labels are missing.

From Descriptions to Training Data

Determine whether a collection of papaya observations is labeled training data.

Inspect the input: Color and softness describe each papaya and form the domain information.

Look for the label: Tastiness is the result associated with each described papaya.

Pair the information: When each description is paired with its tastiness result, the item is a labeled training example.

Check the whole collection: A finite sequence of these labeled pairs is training data, also called a training set.

Descriptions alone are an unlabeled collection of inputs; descriptions paired with tastiness results form labeled training data.

Mistakes About Generalization

  • Treating the lowest training error as proof of good performance over the underlying data distribution.

    A predictor can succeed on the examples it has seen while performing poorly over the underlying distribution. That is overfitting.

    Fix: State separately how the predictor performs on the training sample and how it performs over the underlying data distribution.

  • Defining a hypothesis class as the predictor selected by ERM.

    H is the set of permitted predictors. ERM selects a member h of H after examining the training sample.

    Fix: Use H for the class and h for an individual predictor that belongs to H.

  • Saying that H is chosen after ERM sees the training sample while still calling it prior inductive bias.

    The inductive-bias explanation depends on choosing the hypothesis class before seeing the training data.

    Fix: Describe the order explicitly: choose H, receive S, then apply ERM inside H.

  • Calling an unlabeled collection of inputs training data.

    Training data consists of labeled examples, and each example contains an input from X paired with a label from Y.

    Fix: Check that every input has an associated label before calling the collection labeled training data.

Check Your Understanding

MEDIUM

A learner chooses a hypothesis class H before receiving a finite sequence S of labeled examples. Explain what ERM does after S arrives, why the choice of H is an inductive bias, and why a low error on S does not by itself establish good performance over the underlying data distribution.

Hints
  • Mention that ERM chooses among members of H rather than among every available predictor.
  • Describe S as a sequence of pairs, with an input from X and a label from Y.
  • Use the term overfitting for the case where training performance is good but distribution performance is poor.
EASY

Classify each collection as labeled training data or an unlabeled collection of inputs: a sequence of pairs containing papaya descriptions and tastiness results; a sequence containing papaya descriptions only. Explain the difference using X and Y.

Hints
  • The description is the input from X.
  • The tastiness result is the label from Y.
  • A training example must contain both parts.

The Generalization Picture

  1. Training data is a finite sequence S of labeled pairs, and each pair contains an input from X and a label from Y.
  2. The individual pairs are training examples, while the complete sequence S is also called the training set.
  3. ERM selects a predictor with the lowest training error, but low training error does not guarantee good performance over the underlying data distribution.
  4. Overfitting occurs when a predictor succeeds on the observed training examples but performs poorly over the underlying distribution.
  5. A hypothesis class H restricts ERM to a chosen set of predictors, and choosing H before seeing the training data creates inductive bias.

Key Takeaways

  • A finite sequence of labeled pairs is training data; each pair combines an input from X with a label from Y.
  • ERM minimizes error on the training sample, not automatically over the underlying data distribution.
  • Overfitting is the gap between strong performance on observed examples and poor performance over the underlying distribution.
  • A hypothesis class H is a set of permitted predictors, and ERM chooses within that set.
  • Choosing H before observing the sample supplies inductive bias.