True Error
ERM selects a predictor by minimizing its error on the training sample.
The Learner's Information Problem
A learner wants a prediction rule that makes few errors on the labeled points it may encounter. The difficulty is that the learner does not directly know the entire probability distribution D that produces examples, and it does not directly inspect the unknown target function that supplies their labels. Instead, the learning algorithm receives a training set S. That sample is the evidence available for comparing prediction rules.
Two Sources for Measuring Error
The same prediction rule h can be evaluated in two related ways. True error, also called risk, asks how likely h is to make an error on labeled points drawn according to the probability distribution D. Empirical risk evaluates h using the training data S. The prediction rule is the object being evaluated; D and S are the different sources used for that evaluation.
True error, or risk, is the error of a prediction rule h evaluated with respect to labeled points drawn according to the probability distribution D.
Why Training Data Becomes the Substitute
The learner cannot directly minimize true error because D is unknown. It does, however, have access to S and can inspect how a prediction rule performs on the training examples. Empirical risk is therefore the training-data measure used to compare prediction rules. This does not change the learner's goal: the desired outcome remains a prediction rule with low true error.
Empirical risk is the prediction error of a rule h evaluated using the available training data S.
ERM Selection on One Sample
Comparing Two Candidate Rules
Suppose the same training set S is used to evaluate two candidate prediction rules, h1 and h2. The training observations show fewer errors for h1 than for h2.
Use the same evidence: Evaluate both h1 and h2 on the same training set S. The training set is the evidence supplied to the learner.
Measure each rule: Determine the prediction errors made by h1 on S and the prediction errors made by h2 on S. These produce the empirical-risk measures for the two rules.
Apply the ERM criterion: Because h1 has fewer measured errors on S in this example, ERM favors h1 over h2.
Keep the distinction visible: This selection is based on empirical risk. It does not mean that the learner directly measured either rule's true error with respect to the unknown distribution D.
ERM selects the candidate with the minimum measured error on the training sample; in this generated example, that candidate is h1.
Labels and Prediction Rules
A training set is sampled from an unknown distribution and labeled by an unknown target function. A prediction rule h makes predictions for the examples. Its error is determined by comparing those predictions with the labels carried by the examples. The learner can use these observed comparisons on S, but it cannot directly inspect the complete target function or the complete source distribution.
Common Confusions
Treating training error as true error
Training error is observable on S, whereas true error is defined with respect to labeled points drawn according to the unknown distribution D.
Fix:
Call the training-data measure empirical risk and reserve true error, or risk, for evaluation with respect to D.Saying that ERM directly inspects D
The learner does not know D and cannot directly use knowledge of it.
Fix:
Explain that ERM uses the accessible training set S to compare predictors.Confusing the target function with the learned predictor
The source distinguishes the unknown target function that labels the sample from the predictor produced by the learning algorithm.
Fix:
Keep the roles separate: the target function supplies labels, while h is the prediction rule being evaluated.Claiming that every learning algorithm performs the same candidate search
The ERM criterion concerns which measured error is favored; it does not claim that every algorithm follows one identical search procedure.
Fix:
State the criterion without adding a specific search mechanism: predictors are judged by their error on the training sample, and ERM favors the minimum.
Check Your Understanding
A learning algorithm receives a training set S and compares two prediction rules, hA and hB. hA has the smaller measured error on S. Explain which rule ERM favors, identify whether this comparison uses true error or empirical risk, and state what additional quantity the learner would ideally like to minimize but cannot directly access.
Hints
- Start with the source used for the comparison: S or D.
- ERM favors the candidate with the minimum measured error on the training sample.
- The desired but unavailable measure is the rule's error with respect to the unknown distribution D.
What do you think happens?
A learner can calculate a rule's error on S but does not know D. Which measure can it directly use to compare candidate rules?
Reveal answer
Answer: Empirical risk
Empirical risk evaluates prediction error using the training data S, which the learner has. True error is evaluated with respect to the unknown distribution D.
Key Takeaways
- True error, or risk, evaluates a prediction rule h with respect to labeled points drawn according to the probability distribution D.
- Empirical risk evaluates the same kind of prediction rule using the available training data S.
- The learner wants low true error but cannot directly inspect the unknown distribution D or target function.
- ERM uses the training set as evidence, compares candidate predictors by their measured errors on S, and favors the predictor with the minimum empirical risk.
- Training error and true error are related measures, but they must not be treated as the same quantity.
Key Takeaways
- True error measures how a prediction rule performs with respect to labeled points from the unknown distribution D.
- Empirical risk measures prediction error using the training set S available to the learner.
- The learner wants to minimize true error but directly accesses only training-data evidence.
- Empirical Risk Minimization selects or favors the predictor with the smallest measured error on the training sample.
- The target function supplies labels, while the learned prediction rule h is evaluated against those labels.