Concepts / Estimation Error

Estimation Error

A hypothesis class is the set of possible classifiers from which the learning algorithm chooses.

  • Programming

Choosing Among Classifiers

When a learning algorithm produces a poor predictor, the poor result alone does not tell you why it failed. The model may be too limited to represent the underlying relationship, or it may be so complex that it has fitted noise in the available data. Estimation error is part of the second problem: it concerns the difficulty of accurately estimating the model's true parameters from finite data. To diagnose model selection correctly, first ask whether the chosen hypothesis class is too simple or too rich.

A hypothesis class is the set of possible classifiers from which the learning algorithm chooses.

Class Complexity and Containment

Imagine comparing a small hypothesis class with a richer one. The richer class contains more possible classifiers, so it has more flexibility when representing the real classification problem. A small class offers fewer choices. That restriction can make learning more manageable from a finite sample, but it can also prevent the class from containing a good or optimal classifier.

contained withinRich hypothesisclassMore possible classifiersSmall hypothesisclassFewer possible classifiers
What contains what when comparing a simple hypothesis class with a richer hypothesis class?

Containment is useful because it makes the two pressures visible. A richer class gives the algorithm more candidates and is less likely to block a good solution. A smaller class gives the algorithm fewer candidates to distinguish among, which can reduce estimation error, but the class itself may impose a larger limitation on what can be represented.

Tracing the Two Errors

Approximation error is caused by the hypothesis class itself. It appears when the class is not rich enough to contain the optimal classifier. Estimation error has a different source: the learner must use a finite sample to choose and estimate a model, and the difficulty of doing that is related to both sample size and hypothesis-class complexity.

limits representationlimits choicesadds flexibilityadds choicesSmall classFew classifier choicesApproximation errorMay be higherApproximation errorMay be lowerEstimation errorMay be lowerRich classMany classifier choicesEstimation errorMay be higher
What happens to approximation error and estimation error when the hypothesis class becomes more complex?
Hypothesis-class choiceApproximation errorEstimation error
Very small classMay increase because the optimal classifier may not fitMay decrease because there are fewer choices
Richer classMay decrease because more classifiers are availableMay increase because finite data may not support reliable selection

The two errors respond in opposite directions when class complexity changes.

Increasing richness does not guarantee lower total error. It removes one possible limitation—the class being unable to express a good classifier—but creates more ways to fit the particular sample. With insufficient data, that extra flexibility can increase estimation error. Moving toward a smaller class reverses the pressure: estimation error may decrease while approximation error may increase.

Underfitting and Overfitting

Diagnosing two poor predictors

Two learning systems produce poor predictors. System A uses a very small hypothesis class. System B uses a very rich class and closely follows irregularities in its sample. What kind of limitation should you investigate in each system?

Inspect System A: Because the class is very small, it may be unable to contain the optimal classifier. Its main risk is high approximation error.

Name the behavior: A model that is too simple is associated with underfitting.

Inspect System B: Because the class is very rich, it has many ways to fit the particular sample. With finite data, its estimation may be unreliable.

Name the behavior: A model that is too complex and fits noise is associated with overfitting.

Compare the diagnoses: The two poor results have different likely causes: System A is limited by what it can represent, while System B is at risk of being too closely tied to sample-specific noise.

System A illustrates underfitting and a possible approximation-error problem. System B illustrates overfitting and a possible estimation-error problem.

cannot represent enoughassociated withfinite data and complexityassociated withToo-simple modelLimited representationApproximation errorMay be highUnderfittingModel cannot capture enoughEstimation errorMay be highOverfittingModel fits noiseToo-rich modelMany fitting options
How does a model that is too simple differ from one that is too rich in its fit to training data and true patterns?

Balancing the Selection

Model selection seeks a hypothesis class whose approximation error is not excessively high while its estimation error remains reasonable. A class that is too small may fail to express the real classification problem. A class that is too rich may offer so many alternatives that finite data cannot support reliable estimation. The goal is therefore not maximum simplicity or maximum richness, but a useful balance between the two sources of error.

add useful flexibilityadd unnecessary alternativesRestricted classApproximation error toohighBalanced classBoth errors reasonableVery rich classEstimation error may behigh
How does model selection identify a hypothesis class that balances being expressive enough with being reliable from limited data?

In theory, the ideal hypothesis class would contain only the Bayes optimal classifier. That would avoid choosing among unnecessary alternatives. In practice, this class is generally unavailable because the Bayes optimal classifier depends on the underlying distribution, and that distribution is unknown. If the distribution were already known, there would be little need for learning. This is why practical model selection relies on a reasonable balance rather than a theoretically perfect class.

Using Prior Knowledge

Prior knowledge about the problem can guide the design of a hypothesis class. The knowledge does not need to provide a complete description of the optimal classifier. Even a reasonable conjecture about the structure of the classification problem can help produce a class with acceptable approximation and estimation errors.

The source discussion uses a rectangle to illustrate this principle, not as a universally correct shape for every classification problem. Its role is to show that domain knowledge can motivate a particular class when the optimal classifier is unknown. A useful class is one whose structure is plausible for the problem while still keeping estimation error reasonable.

Common Diagnostic Mistakes

  • Assuming the richest hypothesis class must be best.

    Richness can reduce approximation error, but finite data may make estimation error larger.

    Fix: Look for a class that is expressive enough while keeping estimation error reasonable.

  • Treating approximation error and estimation error as the same limitation.

    A small class may perform poorly because it cannot contain the optimal classifier, which is an approximation-error limitation.

    Fix: Separate the class's representational limitation from the difficulty of estimating a model from finite data.

  • Calling every poor predictor overfitting.

    A poor predictor may be underfitting because it is too simple, or overfitting because it is too complex and fits noise.

    Fix: First determine whether the model is too limited or too closely tied to noise.

  • Assuming that the theoretically ideal class is available in practice.

    The Bayes optimal classifier depends on an unknown distribution.

    Fix: Use prior knowledge and seek a practical balance between the two errors.

Practice Diagnosis

MEDIUM

A learning algorithm uses a class with very few possible classifiers and produces a poor predictor. Explain which error may be high and whether the behavior is associated with underfitting or overfitting. Then describe what risk appears if you replace it with a very rich class while keeping the available sample finite.

Hints
  • Ask whether the small class can contain the optimal classifier.
  • Then ask what many additional classifier choices do to estimation from finite data.
  • Use the terms approximation error, estimation error, underfitting, and overfitting.

What do you think happens?

If a hypothesis class becomes richer, what is the expected pressure on approximation error and estimation error?

  • Both necessarily decrease
  • Approximation error may decrease while estimation error may increase
  • Approximation error may increase while estimation error necessarily decreases
  • Neither error is affected
Reveal answer

Answer: Approximation error may decrease while estimation error may increase.

A richer class is less likely to block a good classifier, but its additional flexibility creates more ways to fit a finite sample. The resulting estimation error may therefore increase.

Key Takeaways

  1. A hypothesis class is the set of classifiers from which the learning algorithm chooses.
  2. Approximation error comes from a class that is not rich enough to contain the optimal classifier.
  3. Estimation error is related to finite sample size and hypothesis-class complexity.
  4. A small class may cause underfitting and higher approximation error, while a very rich class may cause overfitting and higher estimation error.
  5. Good model selection balances expressive power with reliable estimation and uses prior knowledge when available.

Key Takeaways

  • Approximation error measures a limitation in what the chosen hypothesis class can capture.
  • Estimation error is connected to finite data and the complexity of the hypothesis class.
  • Increasing class richness may lower approximation error but may raise estimation error.
  • Underfitting is associated with models that are too simple; overfitting is associated with models that are too complex and fit noise.
  • Model selection seeks a practical balance, guided where possible by prior knowledge about the problem.