Concepts / Relaxing the Realizability Assumption

Relaxing the Realizability Assumption

True error evaluates a prediction rule h with respect to labeled points drawn according to the probability distribution D.

  • Programming

The Evaluation Problem

A learner must choose a prediction rule h, but choosing a rule is not enough. The learner also needs a way to evaluate how well that rule predicts labels. Two evaluation sources matter: the probability distribution D and the training data S. True error uses labeled points drawn according to D. Empirical risk uses the particular training data S available to the learner.

The learner wants a prediction rule with low true error, but the learner does not know D. The learner does have access to S, so empirical risk is the directly available evaluation measure.

What do you think happens?

A learner can inspect the training data S but does not know the distribution D. Which measure can the learner calculate directly from the information it has?

  • True error
  • Empirical risk
  • Both measures equally directly
Reveal answer

Answer: Empirical risk

Empirical risk evaluates prediction error using S, which the learner has. True error is defined using labeled points drawn according to D, but D is unknown to the learner.

True Error Under D

The true error, also called the risk, of a prediction rule h asks how likely h is to make an error on labeled points drawn according to the probability distribution D. The prediction rule is evaluated against the labeled examples that D produces, rather than only against the examples that happened to appear in one training sample.

The distribution D matters in two connected ways. It determines which labeled points are relevant to the evaluation, and it determines how those points contribute to the overall assessment through the distribution from which they are drawn. True error therefore describes performance with respect to D, not merely performance on a finite list of observed examples.

drawsprovides labeled casesis compared onaggregatesDistribution Dsource of labeled pointsLabeled pointsdrawn according to DPrediction errorsevaluated over DTrue errorrisk of hPrediction rule hmakes predictions
How does the distribution D determine which labeled examples count, and how much they contribute, when measuring the average error of a prediction rule h?

Evaluating the Same Rule Against D

Consider a prediction rule h that is evaluated on labeled points drawn according to a distribution D.

Identify the prediction rule: The object being evaluated is h. True error is not a property of the training data alone; it evaluates this prediction rule.

Use labeled points associated with D: The relevant labeled points are drawn according to D. The evaluation is therefore tied to the distribution that produces the points.

Assess the errors: For each evaluated point, h either makes an error or does not. True error describes how likely an error is for points drawn according to D.

True error is the risk of h with respect to D: it describes the error behavior of h on labeled points drawn according to that distribution.

Training Data and Empirical Risk

Empirical risk evaluates the prediction error of h using the training data S. It is an evaluation based on the finite training sample available to the learner.

Empirical risk and true error evaluate the same kind of object: a prediction rule h. Their difference is the source used for evaluation. True error uses labeled points associated with the unknown distribution D. Empirical risk uses the observed training data S.

providesevaluated by hcontainsaggregatedDistribution Dunknown to learnerLabeled pointsdrawn according to DTrue errorrisk of hTraining data Savailable to learnerObserved errorsmade by h on SEmpirical riskerror on S
What is the difference between the error h makes on the entire distribution D and the average error h makes on the finite training sample S?
MeasureEvaluation sourceLearner's access
True errorLabeled points drawn according to DD is unknown
Empirical riskTraining data SS is available

Both measures evaluate prediction errors made by h, but they use different evaluation sources.

From Observable Risk to Desired Risk

The learner's goal and the learner's information are not the same thing. The learner wants a prediction rule h with low true error because true error describes performance with respect to D. However, the learner cannot directly use D because D is unknown. The learner can inspect S and use empirical risk as an accessible way to evaluate h.

evaluates h oninformsaims for lowdefines evaluationTraining data SobservableEmpirical riskcomputed from SChoose hlearner's decisionTrue errordesired evaluation under DDistribution Dunknown to learner
How does the learner use the observable training data S and empirical risk to approximate the unobservable true error under D?

When Perfect Prediction Is Not Required

Relaxing the realizability assumption means that the discussion does not require a prediction rule to classify every possible example associated with D correctly. Some errors may remain part of the evaluation. In that setting, the useful question is not whether h makes no errors everywhere, but how much error h has under D and how much error it has on the observed training data S.

evaluated under strict requirementevaluated with Devaluated with SPrediction rule hperfect classificationexpectedNo errorsunder the requirementPrediction rule hsome errors allowedTrue errorerror under DEmpirical riskerror on S
What changes when no prediction rule is required to classify every possible example from D correctly, and how do true error and empirical risk reflect those mistakes?

This change does not replace the two evaluation measures. True error still uses D, and empirical risk still uses S. Relaxing the assumption changes how the errors are interpreted: errors are quantities to evaluate and compare rather than evidence that the entire learning setup has failed simply because a perfect rule is unavailable.

Common Misunderstandings

  • Treating true error as the error measured only on S.

    True error is evaluated with respect to labeled points drawn according to D, while S is the source for empirical risk.

    Fix: Use true error for the distribution-based evaluation and empirical risk for the training-data evaluation.

  • Assuming the learner knows D.

    The learner does not know the data-generating distribution D.

    Fix: Remember that the learner has access to S and can evaluate empirical risk from that training data.

  • Assuming that the learner's accessible measure is automatically the learner's final goal.

    The learner wants a prediction rule with low true error, but uses information from S because D is unknown.

    Fix: Keep the goal, low true error, separate from the accessible evaluation, empirical risk.

  • Interpreting relaxed realizability as ignoring errors.

    True error and empirical risk still measure the errors made by h; the distinction between D and S remains central.

    Fix: Evaluate the remaining errors using the appropriate source: D for true error and S for empirical risk.

Check Your Understanding

EASY

For a prediction rule h, write a short explanation of the difference between its true error and its empirical risk. Your explanation must name the evaluation source for each measure and state which source the learner can directly access.

Hints
  • Start by identifying what D supplies for the true-error evaluation.
  • Then identify what S supplies for the empirical-risk evaluation.
  • End by separating the learner's desired objective from the information available during learning.

A Complete Distinction

Explain what is being measured when a learner evaluates h using D and when the learner evaluates h using S.

Evaluation with D: This measures the true error, or risk, of h on labeled points drawn according to the probability distribution D.

Evaluation with S: This measures empirical risk by using the training data S available to the learner.

Learning implication: The learner wants low true error but cannot directly use D because D is unknown. The learner can use empirical risk because S is available.

True error is the desired distribution-based evaluation; empirical risk is the training-data-based evaluation available to the learner.

Key Takeaways

  1. True error, or risk, evaluates a prediction rule h on labeled points drawn according to the probability distribution D.
  2. Empirical risk evaluates the same prediction rule h using the training data S.
  3. The learner wants low true error but does not know D.
  4. The learner does have access to S, so empirical risk is the directly available evaluation measure.
  5. Relaxing the realizability assumption allows the discussion to focus on measuring and comparing errors even when perfect classification is not required.

Key Takeaways

  • True error measures how likely a prediction rule h is to make an error on labeled points drawn according to D.
  • Empirical risk measures h's error on the training data S.
  • The learner wants to minimize true error but cannot directly access D.
  • Because S is available, the learner can directly evaluate empirical risk.
  • Relaxing realizability means that evaluation can account for prediction errors without requiring a perfectly correct rule.