Training Error
ERM selects a predictor by minimizing its error on the training sample.
The Learner’s Available Evidence
A learning algorithm must produce a predictor from a training set. The training set is sampled from an unknown distribution and labeled by an unknown target function. Because the learner cannot directly inspect the entire source distribution or the target function, it cannot directly measure the predictor’s error in the complete data-generating situation. It can, however, inspect the training examples it was given.
Training error is the error made by a classifier over the particular training sample available to the learner.
What Training Error Measures
Training error measures the error made by a classifier on its training sample. The phrase on its training sample is essential: this is a sample-based measurement, not a statement about the classifier’s error everywhere.
The learner has access to the training examples and their labels, so it can check how the classifier performs on those examples. This makes training error observable and calculable. The broader error that would be measured with respect to the unknown data distribution and unknown target function is not directly available to the learner.
| Measurement | What it uses | What the learner can do |
|---|---|---|
| Training error | The classifier’s errors over the training sample | Calculate it from available information |
| True error | The broader data-generating situation involving the unknown distribution and target function | It is not directly available |
Selecting a Predictor by Sample Error
Empirical Risk Minimization, or ERM, uses the training sample as evidence for comparing predictors. The learner evaluates the errors that candidate predictors incur on that same sample and favors the predictor with the minimum measured error. This describes the selection criterion; it does not claim that every learning algorithm searches through candidates by one identical mechanical procedure.
Comparing Three Candidate Predictors
A learner evaluates three candidate predictors on the same training sample. Predictor A makes 4 measured errors, Predictor B makes 1 measured error, and Predictor C makes 3 measured errors. Which predictor does ERM favor?
Measure: The learner evaluates each candidate on the available training sample.
Compare: The measured errors are compared using the same sample as evidence for every candidate.
Select: Predictor B has the minimum measured error among the three candidates.
ERM favors Predictor B because it incurs the fewest errors on the training sample.
Sample Error Versus Broader Error
Training error and true error answer different questions. Training error asks how many errors the classifier incurs on the particular training sample. True error concerns the broader situation associated with the unknown data distribution and target function. Since the learner cannot directly inspect that full situation, it cannot directly minimize true error. ERM instead uses the measurable sample-based error as its selection criterion.
What do you think happens?
A classifier has the fewest measured errors on the training sample. What can the learner conclude immediately?
Reveal answer
Answer: That the classifier has the fewest measured errors on that training sample
Training error is tied to the supplied sample. The learner can use it to compare predictors on that evidence, but true error involving the unknown distribution and target function is not directly available.
Names for the Same Sample-Based Idea
You may encounter empirical error and empirical risk used interchangeably with training error. In this context, these terms refer to the error that a classifier incurs over the training sample. The terminology does not change the practical question: how many errors does the classifier make on the finite sample available to the learner?
Mistakes About What It Proves
Treating training error as the classifier’s error everywhere.
Training error is specifically the error over the training sample. It is not directly available information about the full unknown distribution and target function.
Fix:
State the conclusion narrowly: the classifier made few measured mistakes on that training sample.Saying that ERM directly minimizes true error.
The learner cannot directly inspect the full source distribution or the target function.
Fix:
Say that ERM selects a predictor by minimizing its measured error on the training sample.Treating the training set as irrelevant after the predictor is selected.
A predictor produced through learning depends on the training set supplied to the algorithm.
Fix:
Remember that the training set is the evidence used to compare predictors and influences the predictor produced.Assuming training error, empirical error, and empirical risk are three unrelated measurements in this context.
The source uses empirical error and empirical risk as terms often interchangeable with training error.
Fix:
Treat them here as related names for the classifier’s error over the training sample.
Check Your Understanding
Explain in your own words why a learner can calculate training error but cannot directly calculate true error. Then describe how ERM uses training error to choose between two candidate predictors.
Hints
- Begin with the information contained in the training sample.
- Mention the unknown data distribution and target function.
- End by identifying which candidate has the minimum measured error on the sample.
Classifying the Measurement
A statement says: A classifier makes two mistakes on the examples supplied to the learning algorithm. Is this a statement about training error or true error?
Locate the evidence: The statement refers to examples supplied to the learning algorithm, so it refers to the training sample.
Identify the measurement: The classifier’s mistakes over that training sample are the training error.
Check the scope: The statement does not describe the classifier’s error over the unknown data distribution and target function.
The statement describes training error, not directly available true error.
Key Takeaways
- Training error is the error a classifier incurs over its training sample.
- The learner can calculate training error because the training sample is available to it.
- True error concerns the unknown data distribution and target function, so it is not directly available to the learner.
- ERM selects a predictor by comparing candidate predictors using their errors on the training sample and favoring the minimum measured error.
- Empirical error and empirical risk are often interchangeable terms for training error in this context.
Key Takeaways
- Training error measures classifier mistakes on the available training sample.
- A learner uses training error because it cannot directly inspect the unknown data distribution or target function.
- Empirical Risk Minimization compares candidate predictors on the same training sample and favors the one with minimum measured error.
- A low training error describes performance on the observed sample; it does not directly reveal the classifier’s broader true error.
- Empirical error and empirical risk are often used as interchangeable names for training error.