Concepts / Empirical Risk

Empirical Risk

Approximation error measures the gap between the best hypothesis in a class and the true hypothesis.

  • Programming

The Evidence a Learner Can See

A learning algorithm receives a training set sampled from an unknown distribution and labeled by an unknown target function. It must choose a predictor without directly inspecting the entire source distribution or the target function. The error measured on the available training sample is observable to the learner; the error over the unknown distribution and target function is not directly available.

Empirical risk is the error a predictor incurs on the training sample. It gives the learner a concrete basis for comparing predictors, but it is not automatically the same as the predictor's true risk outside that sample.

sampleslabelsinforms learningmeasured on sampleevaluated over unknown sourceUnknowndistributionTraining samplefinite evidencePredictorselected from evidenceEmpirical riskobserved on training sampleUnknown targetfunctionTrue riskover distribution andtarget
How does a finite training sample provide an observed training error that may differ from the unknown error over the true data distribution and target function?

Two Sources of Error

A poor learned result can have more than one explanation. The hypothesis class may be unable to represent the underlying relationship well. Alternatively, the class may be capable of doing well, but the learning process may select a poor hypothesis because it learned from only a finite sample. These possibilities are described by approximation error and estimation error.

Approximation error is the gap between the true hypothesis and the best hypothesis available in a chosen hypothesis class. It concerns what the class can achieve in principle.

Estimation error is the difference between empirical risk and true risk caused by learning from a finite sample. It concerns the gap created when the learner infers a particular hypothesis from limited evidence.

gap to best class memberbest available in classreference predictorselected from finite sampleTrue hypothesisApproximation errorclass limitationBest class memberEstimation errorfinite-sample learningBest class memberLearned hypothesis
How do the errors caused by an overly limited hypothesis class differ from the errors caused by learning from a finite sample?

How ERM Chooses a Predictor

Empirical Risk Minimization, or ERM, selects a predictor by minimizing its error on the training sample.

The learner cannot directly minimize true risk because it cannot directly inspect the entire unknown distribution or the unknown target function. Instead, it uses the training set as evidence. Candidate predictors are compared according to the errors they incur on that sample, and ERM favors the predictor with the smallest measured error.

provides evidenceevaluaterank errorsfavorTraining sampleCandidate predictorsCompare sample errorsempirical risksSmallest empiricalriskSelected predictor
How does a learner compare candidate predictors on a training sample and select the one with the smallest empirical risk?

Comparing Two Candidate Predictors

A training sample is used to compare Predictor A and Predictor B. Predictor A makes fewer errors on the training sample than Predictor B. Which predictor does ERM favor?

Measure: Evaluate the errors made by both predictors on the training sample.

Compare: Place the measured empirical risks side by side.

Select: ERM favors the predictor with the smaller measured training error.

ERM favors Predictor A. This conclusion concerns empirical risk on the training sample; it does not by itself establish which predictor has lower true risk.

What do you think happens?

Predictor A has a smaller training error than Predictor B. Which predictor does ERM favor?

  • Predictor A
  • Predictor B
  • Neither, because ERM uses validation error
Reveal answer

Answer: Predictor A

ERM uses the error measured on the training sample and favors the predictor with the minimum measured error. This choice does not by itself prove that Predictor A has lower true risk.

Validation Beyond the Training Sample

Validation evaluates a learned hypothesis on data separate from the training sample. Because the hypothesis was not fitted using those validation examples, validation provides evidence about how the learned hypothesis performs outside the sample used to fit it. The source describes validation as a way to estimate the true risk of the learned hypothesis, with the difference between distribution risk and validation risk capable of being bounded quite tightly using Theorem 11.1.

fitsevaluates onprovides separate evidenceTraining sampleused to fitLearned hypothesisValidation riskestimateevidence about learnedhypothesisValidation setseparate data
How does evaluating a hypothesis on data separate from the training sample provide evidence about how well it generalizes?

Reading Risk Patterns

Training and validation or true risk together provide more information than training risk alone. Large empirical risk indicates underfitting: the predictor performs poorly even on the training sample. A large gap between empirical risk and true risk indicates overfitting: the predictor appears better on the training sample than it does outside that sample.

can be associated withcan be associated withlarge valuesmall relative valuelarge gap from training riskLow complexityUnderfittinglarge empirical riskTraining riskobserved on sampleHigher complexityOverfittinglargeempirical-to-true-risk gapValidation or trueriskoutside training sample
What happens to training risk and validation or true risk as model complexity increases, and where do underfitting and overfitting appear?
Observed patternPlausible issueInterpretation
Large empirical riskUnderfittingThe predictor performs poorly on the training sample.
Small empirical risk with a large empirical-to-true-risk gapOverfittingThe predictor performs much better on the training sample than outside it.
Training risk by itselfInsufficient diagnosisIt does not necessarily reveal approximation error.

Choosing a Remedy

When a model performs poorly, first ask which error source is most plausible. If the hypothesis class is too limited to represent the underlying relationship well, the concern is approximation error and a more expressive or otherwise different class may be worth considering. If the learned hypothesis performs well on training data but poorly outside it, the concern is estimation error or overfitting; more data or stronger regularization may be relevant remedies. These choices are diagnostic directions, not guarantees, because validation does not directly measure approximation error.

inspectinspectmay suggestconsidermay suggestconsiderconsiderPoor resultLarge empirical riskClass limitationapproximation concernMore expressive classLargeempirical-to-true-riskgapFinite-samplelearning issueestimation concernMore dataStrongerregularization
How does identifying approximation error versus estimation error determine whether to use a more expressive model, more data, or stronger regularization?

Common Diagnostic Mistakes

  • Treating training risk as a direct measurement of approximation error.

    Training risk concerns the selected predictor on one finite sample. Approximation error concerns the gap between the true hypothesis and the best member of the entire class.

    Fix: Use training and validation or true-risk evidence to assess the learned hypothesis, while remembering that approximation error remains harder to estimate directly.

  • Assuming validation directly reveals approximation error.

    Validation evaluates the particular learned hypothesis, not necessarily the best possible member of the class.

    Fix: Use validation to assess the learned hypothesis outside the training sample, not as a complete measurement of approximation error.

  • Assuming ERM minimizes true risk.

    The learner cannot directly inspect the unknown distribution or target function.

    Fix: State that ERM minimizes measured error on the training sample.

  • Calling every poor result underfitting.

    Poor performance can arise from a limited class or from learning a poor hypothesis from a finite sample.

    Fix: Compare empirical risk with outside-sample performance to investigate approximation and estimation concerns separately.

Check Your Understanding

MEDIUM

A learning algorithm compares two predictors on a training sample. Predictor R has the smaller training error, but its performance on a separate validation set is much worse than its training performance. Explain which predictor ERM selects, what pattern the validation result suggests, and why the result still does not directly measure approximation error.

Hints
  • ERM selects using error measured on the training sample.
  • Compare empirical risk with performance outside the training sample.
  • Approximation error concerns the best member of the hypothesis class, not only the selected predictor.

Answering the Diagnostic Question

Predictor R has the smallest training error, but its validation error is much larger.

Selection: ERM selects Predictor R because it has the smallest measured error on the training sample.

Pattern: The large difference between training and validation performance suggests a large empirical-to-true-risk gap and therefore an overfitting or estimation concern.

Limitation: The validation result concerns the learned Predictor R. It does not directly tell us how well the best member of the whole hypothesis class could perform.

ERM selection, overfitting diagnosis, and approximation-error diagnosis are three related but distinct conclusions.

Key Takeaways

  1. Empirical risk is the error measured on a finite training sample and is observable to the learner.
  2. ERM selects a predictor by favoring the smallest training-sample error because true risk is not directly available.
  3. Approximation error concerns the gap between the true hypothesis and the best member of a hypothesis class.
  4. Estimation error concerns the gap between empirical risk and true risk caused by learning from a finite sample.
  5. Large empirical risk suggests underfitting, while a large empirical-to-true-risk gap suggests overfitting.
  6. Validation helps assess the learned hypothesis outside the training sample, but it does not directly measure approximation error.

Key Takeaways

  • Empirical risk measures observed performance on a finite training sample, not automatically performance over the unknown distribution.
  • ERM uses training error because the learner cannot directly access true risk.
  • Approximation error describes a limitation of the hypothesis class, while estimation error comes from learning from finite data.
  • Validation provides evidence about the learned hypothesis outside the training sample but does not directly reveal approximation error.
  • Training and validation patterns help distinguish likely underfitting from likely overfitting and guide more informed remedies.