Concepts / Training Error

Training Error

ERM selects a predictor by minimizing its error on the training sample.

  • Programming

The Learner’s Available Evidence

A learning algorithm must produce a predictor from a training set. The training set is sampled from an unknown distribution and labeled by an unknown target function. Because the learner cannot directly inspect the entire source distribution or the target function, it cannot directly measure the predictor’s error in the complete data-generating situation. It can, however, inspect the training examples it was given.

Training error is the error made by a classifier over the particular training sample available to the learner.

evaluate on examplescount sample errorswould be neededwould be neededTraining setavailable examplesClassifiercandidate predictorTraining errorcalculableData distributionunknownTarget functionunknownTrue errornot directly available
Which information can the learner inspect to calculate training error, and which information needed for true error is unavailable?

What Training Error Measures

Training error measures the error made by a classifier on its training sample. The phrase on its training sample is essential: this is a sample-based measurement, not a statement about the classifier’s error everywhere.

The learner has access to the training examples and their labels, so it can check how the classifier performs on those examples. This makes training error observable and calculable. The broader error that would be measured with respect to the unknown data distribution and unknown target function is not directly available to the learner.

MeasurementWhat it usesWhat the learner can do
Training errorThe classifier’s errors over the training sampleCalculate it from available information
True errorThe broader data-generating situation involving the unknown distribution and target functionIt is not directly available

Selecting a Predictor by Sample Error

Empirical Risk Minimization, or ERM, uses the training sample as evidence for comparing predictors. The learner evaluates the errors that candidate predictors incur on that same sample and favors the predictor with the minimum measured error. This describes the selection criterion; it does not claim that every learning algorithm searches through candidates by one identical mechanical procedure.

evaluateevaluateevaluateerror on sampleerror on sampleerror on samplechoose minimumTraining samplecommon evidencePredictor Ameasured errorPredictor Bmeasured errorPredictor Cmeasured errorCompare errorson the sampleSelected predictorminimum measured error
How does a learner evaluate several candidate predictors on the same training examples and select the one with the fewest mistakes?

Comparing Three Candidate Predictors

A learner evaluates three candidate predictors on the same training sample. Predictor A makes 4 measured errors, Predictor B makes 1 measured error, and Predictor C makes 3 measured errors. Which predictor does ERM favor?

Measure: The learner evaluates each candidate on the available training sample.

Compare: The measured errors are compared using the same sample as evidence for every candidate.

Select: Predictor B has the minimum measured error among the three candidates.

ERM favors Predictor B because it incurs the fewest errors on the training sample.

Sample Error Versus Broader Error

measure on observed samplewould require unavailable informationTraining examplesobserved sampleSample mistakestraining errorNew examplesunknown distributionBroader mistakestrue error
What is the difference between counting mistakes on the observed training sample and measuring mistakes over new examples from the unknown data distribution?

Training error and true error answer different questions. Training error asks how many errors the classifier incurs on the particular training sample. True error concerns the broader situation associated with the unknown data distribution and target function. Since the learner cannot directly inspect that full situation, it cannot directly minimize true error. ERM instead uses the measurable sample-based error as its selection criterion.

What do you think happens?

A classifier has the fewest measured errors on the training sample. What can the learner conclude immediately?

  • That the classifier has the fewest errors everywhere
  • That the classifier has the fewest measured errors on that training sample
  • That the unknown target function has been fully identified
Reveal answer

Answer: That the classifier has the fewest measured errors on that training sample

Training error is tied to the supplied sample. The learner can use it to compare predictors on that evidence, but true error involving the unknown distribution and target function is not directly available.

Names for the Same Sample-Based Idea

You may encounter empirical error and empirical risk used interchangeably with training error. In this context, these terms refer to the error that a classifier incurs over the training sample. The terminology does not change the practical question: how many errors does the classifier make on the finite sample available to the learner?

describesoften namesoften namesTraining errorclassifier on trainingsampleEmpirical errorsample-based errorEmpirical risksample-based errorError on trainingsampleshared context
How are training error, empirical error, and empirical risk related when describing a predictor’s mistakes on a finite sample?

Mistakes About What It Proves

  • Treating training error as the classifier’s error everywhere.

    Training error is specifically the error over the training sample. It is not directly available information about the full unknown distribution and target function.

    Fix: State the conclusion narrowly: the classifier made few measured mistakes on that training sample.

  • Saying that ERM directly minimizes true error.

    The learner cannot directly inspect the full source distribution or the target function.

    Fix: Say that ERM selects a predictor by minimizing its measured error on the training sample.

  • Treating the training set as irrelevant after the predictor is selected.

    A predictor produced through learning depends on the training set supplied to the algorithm.

    Fix: Remember that the training set is the evidence used to compare predictors and influences the predictor produced.

  • Assuming training error, empirical error, and empirical risk are three unrelated measurements in this context.

    The source uses empirical error and empirical risk as terms often interchangeable with training error.

    Fix: Treat them here as related names for the classifier’s error over the training sample.

Check Your Understanding

MEDIUM

Explain in your own words why a learner can calculate training error but cannot directly calculate true error. Then describe how ERM uses training error to choose between two candidate predictors.

Hints
  • Begin with the information contained in the training sample.
  • Mention the unknown data distribution and target function.
  • End by identifying which candidate has the minimum measured error on the sample.

Classifying the Measurement

A statement says: A classifier makes two mistakes on the examples supplied to the learning algorithm. Is this a statement about training error or true error?

Locate the evidence: The statement refers to examples supplied to the learning algorithm, so it refers to the training sample.

Identify the measurement: The classifier’s mistakes over that training sample are the training error.

Check the scope: The statement does not describe the classifier’s error over the unknown data distribution and target function.

The statement describes training error, not directly available true error.

Key Takeaways

  1. Training error is the error a classifier incurs over its training sample.
  2. The learner can calculate training error because the training sample is available to it.
  3. True error concerns the unknown data distribution and target function, so it is not directly available to the learner.
  4. ERM selects a predictor by comparing candidate predictors using their errors on the training sample and favoring the minimum measured error.
  5. Empirical error and empirical risk are often interchangeable terms for training error in this context.

Key Takeaways

  • Training error measures classifier mistakes on the available training sample.
  • A learner uses training error because it cannot directly inspect the unknown data distribution or target function.
  • Empirical Risk Minimization compares candidate predictors on the same training sample and favors the one with minimum measured error.
  • A low training error describes performance on the observed sample; it does not directly reveal the classifier’s broader true error.
  • Empirical error and empirical risk are often used as interchangeable names for training error.