Relaxing the Realizability Assumption
True error evaluates a prediction rule h with respect to labeled points drawn according to the probability distribution D.
The Evaluation Problem
A learner must choose a prediction rule h, but choosing a rule is not enough. The learner also needs a way to evaluate how well that rule predicts labels. Two evaluation sources matter: the probability distribution D and the training data S. True error uses labeled points drawn according to D. Empirical risk uses the particular training data S available to the learner.
The learner wants a prediction rule with low true error, but the learner does not know D. The learner does have access to S, so empirical risk is the directly available evaluation measure.
What do you think happens?
A learner can inspect the training data S but does not know the distribution D. Which measure can the learner calculate directly from the information it has?
Reveal answer
Answer: Empirical risk
Empirical risk evaluates prediction error using S, which the learner has. True error is defined using labeled points drawn according to D, but D is unknown to the learner.
True Error Under D
The true error, also called the risk, of a prediction rule h asks how likely h is to make an error on labeled points drawn according to the probability distribution D. The prediction rule is evaluated against the labeled examples that D produces, rather than only against the examples that happened to appear in one training sample.
The distribution D matters in two connected ways. It determines which labeled points are relevant to the evaluation, and it determines how those points contribute to the overall assessment through the distribution from which they are drawn. True error therefore describes performance with respect to D, not merely performance on a finite list of observed examples.
Evaluating the Same Rule Against D
Consider a prediction rule h that is evaluated on labeled points drawn according to a distribution D.
Identify the prediction rule: The object being evaluated is h. True error is not a property of the training data alone; it evaluates this prediction rule.
Use labeled points associated with D: The relevant labeled points are drawn according to D. The evaluation is therefore tied to the distribution that produces the points.
Assess the errors: For each evaluated point, h either makes an error or does not. True error describes how likely an error is for points drawn according to D.
True error is the risk of h with respect to D: it describes the error behavior of h on labeled points drawn according to that distribution.
Training Data and Empirical Risk
Empirical risk evaluates the prediction error of h using the training data S. It is an evaluation based on the finite training sample available to the learner.
Empirical risk and true error evaluate the same kind of object: a prediction rule h. Their difference is the source used for evaluation. True error uses labeled points associated with the unknown distribution D. Empirical risk uses the observed training data S.
| Measure | Evaluation source | Learner's access |
|---|---|---|
| True error | Labeled points drawn according to D | D is unknown |
| Empirical risk | Training data S | S is available |
Both measures evaluate prediction errors made by h, but they use different evaluation sources.
From Observable Risk to Desired Risk
The learner's goal and the learner's information are not the same thing. The learner wants a prediction rule h with low true error because true error describes performance with respect to D. However, the learner cannot directly use D because D is unknown. The learner can inspect S and use empirical risk as an accessible way to evaluate h.
When Perfect Prediction Is Not Required
Relaxing the realizability assumption means that the discussion does not require a prediction rule to classify every possible example associated with D correctly. Some errors may remain part of the evaluation. In that setting, the useful question is not whether h makes no errors everywhere, but how much error h has under D and how much error it has on the observed training data S.
This change does not replace the two evaluation measures. True error still uses D, and empirical risk still uses S. Relaxing the assumption changes how the errors are interpreted: errors are quantities to evaluate and compare rather than evidence that the entire learning setup has failed simply because a perfect rule is unavailable.
Common Misunderstandings
Treating true error as the error measured only on S.
True error is evaluated with respect to labeled points drawn according to D, while S is the source for empirical risk.
Fix:
Use true error for the distribution-based evaluation and empirical risk for the training-data evaluation.Assuming the learner knows D.
The learner does not know the data-generating distribution D.
Fix:
Remember that the learner has access to S and can evaluate empirical risk from that training data.Assuming that the learner's accessible measure is automatically the learner's final goal.
The learner wants a prediction rule with low true error, but uses information from S because D is unknown.
Fix:
Keep the goal, low true error, separate from the accessible evaluation, empirical risk.Interpreting relaxed realizability as ignoring errors.
True error and empirical risk still measure the errors made by h; the distinction between D and S remains central.
Fix:
Evaluate the remaining errors using the appropriate source: D for true error and S for empirical risk.
Check Your Understanding
For a prediction rule h, write a short explanation of the difference between its true error and its empirical risk. Your explanation must name the evaluation source for each measure and state which source the learner can directly access.
Hints
- Start by identifying what D supplies for the true-error evaluation.
- Then identify what S supplies for the empirical-risk evaluation.
- End by separating the learner's desired objective from the information available during learning.
A Complete Distinction
Explain what is being measured when a learner evaluates h using D and when the learner evaluates h using S.
Evaluation with D: This measures the true error, or risk, of h on labeled points drawn according to the probability distribution D.
Evaluation with S: This measures empirical risk by using the training data S available to the learner.
Learning implication: The learner wants low true error but cannot directly use D because D is unknown. The learner can use empirical risk because S is available.
True error is the desired distribution-based evaluation; empirical risk is the training-data-based evaluation available to the learner.
Key Takeaways
- True error, or risk, evaluates a prediction rule h on labeled points drawn according to the probability distribution D.
- Empirical risk evaluates the same prediction rule h using the training data S.
- The learner wants low true error but does not know D.
- The learner does have access to S, so empirical risk is the directly available evaluation measure.
- Relaxing the realizability assumption allows the discussion to focus on measuring and comparing errors even when perfect classification is not required.
Key Takeaways
- True error measures how likely a prediction rule h is to make an error on labeled points drawn according to D.
- Empirical risk measures h's error on the training data S.
- The learner wants to minimize true error but cannot directly access D.
- Because S is available, the learner can directly evaluate empirical risk.
- Relaxing realizability means that evaluation can account for prediction errors without requiring a perfectly correct rule.