Concepts / Classification Error

Classification Error

The Bayes optimal predictor f_D is defined relative to a probability distribution D.

  • Programming

The Best Possible Benchmark

A classifier receives an input and returns one of two labels, 0 or 1. Its classification error is determined by how often its predicted label differs from the true label on examples produced by a distribution D. The central question is not always whether a classifier can avoid every mistake. A more precise question is: among all classifiers, which one has the lowest error for the distribution that generates the data?

The Bayes optimal predictor, written f_D, is the classifier selected as the best possible label-predicting function for a particular distribution D. It maps inputs from X to binary labels in {0, 1}. The subscript D is essential: this predictor is defined relative to the distribution being considered, so changing D can change which classifier is optimal.

Tracing a Single Input

Fix one input x from X. The distribution D describes how inputs occur together with labels. When x appears, the two possible labels are 0 and 1. The Bayes optimal predictor must choose one of them. Its choice is not an independent rule that works without context; it is the choice that belongs to the complete function that achieves the lowest error for D. Viewed point by point, the relevant comparison is which label is more likely for that input under D.

D relates x to 0D relates x to 1lower-error choicelower-error choiceinput xfrom Xlabel 0likelihood under Df_D(x)selected binary labellabel 1likelihood under D
For a fixed input x, how does comparing the two possible labels support the Bayes optimal choice?

One Input, Two Possible Labels

Consider a fixed distribution D and one input x. A classifier must return either 0 or 1.

Identify the available predictions: The predictor can assign x only one of the two binary labels: 0 or 1.

Use the distribution: The choice must be evaluated according to how D relates this input to the two possible labels.

Connect the local choice to the full predictor: The selected label is not considered in isolation. It is part of the complete function f_D, which is defined to have the lowest error under D.

The Bayes optimal prediction for x is the label choice that contributes to the lowest-error classifier for distribution D.

How Distribution D Determines the Predictor

The distribution D is not merely background information. It determines which inputs and labels occur together and therefore determines what counts as a good prediction. Once D is fixed, f_D is defined as the best function from X to {0, 1} for that distribution. The same input space can therefore be associated with different Bayes optimal predictors when the underlying distribution changes.

describes occurrencesprovides label relationshipis evaluateddefines predictiondistribution Dinputs with binary labelsinput xfrom Xf_Dbest function for Dlabel comparisonunder D
How does the input-label distribution D determine the Bayes optimal prediction for each input?

The notation f_D reminds you that the predictor cannot be defined independently of the distribution. D describes the input-label behavior; f_D is the best classifier for that behavior.

Why Its Error Cannot Be Lowered

Take any classifier g that maps X to {0, 1}. Under the same distribution D, compare the mistakes made by g with the mistakes made by f_D. By definition, f_D is the classifier with the lowest error available for D. Therefore, the error of f_D is less than or equal to the error of g. This is the optimality guarantee.

evaluatesevaluateshashasno greater thandistribution Dsame inputs and labelsf_DBayes optimal predictorBayes errorlowest error under Dgarbitrary classifiererror of gat least Bayes error
How do the errors of the Bayes optimal predictor and an arbitrary classifier compare on the same distribution?

Comparing Two Classifiers

For one fixed distribution D, compare the Bayes optimal predictor f_D with another classifier g.

Hold D fixed: Both classifiers must be evaluated on the same distribution of inputs and binary labels.

Evaluate f_D: Because f_D is defined as the best classifier for D, its error is the minimum achievable error under that distribution.

Evaluate g: The arbitrary classifier g may match that minimum, or it may make more mistakes. It cannot have a lower error than f_D under D.

The Bayes optimal predictor has error no greater than the error of the arbitrary classifier.

The Practical Knowledge Gap

The Bayes optimal predictor is an ideal benchmark because it is defined using D. In practice, the obstacle identified here is that D is unknown. Without knowing the distribution that generates the input-label pairs, we cannot directly construct or utilize the predictor that is optimal for that distribution.

needed to definemissing informationcannot be directly utilizedunknown Ddata-generatingdistributionf_Ddefined relative to Ddirect useunavailableD is not known
What information must connect the data distribution to the ideal predictor, and what happens when that distribution is unknown?

Mistakes to Avoid

  • Treating Bayes optimal as meaning zero error.

    Optimality means lowest error under D, not necessarily no mistakes.

    Fix: Allow for the possibility that the input-label relationship is not perfectly predictable.

  • Ignoring the subscript D.

    The predictor is defined relative to a probability distribution, and the best function depends on that distribution.

    Fix: Read f_D as the predictor that is optimal for this particular D.

  • Comparing classifiers under different distributions.

    The optimality guarantee compares errors under the same distribution D.

    Fix: Hold D fixed before comparing the error of f_D with the error of any classifier g.

  • Assuming the ideal predictor can be used directly when D is unknown.

    The source identifies unknown D as the practical obstacle to directly utilizing f_D.

    Fix: Treat f_D as an ideal benchmark unless the relevant distribution is known.

Check Your Understanding

MEDIUM

A classifier g and the Bayes optimal predictor f_D are evaluated under the same distribution D. State the guaranteed relationship between their errors, then explain whether f_D must have zero error. Finally, explain why knowing the identity of f_D does not automatically make it usable when D is unknown.

Hints
  • Use the definition of optimality: compare f_D with every classifier under the same D.
  • Separate lowest possible error from zero error.
  • Ask what information is required to define f_D.
  1. The Bayes optimal predictor f_D is the best function from X to the binary labels 0 and 1 for a fixed distribution D. Its error is no greater than the error of any other classifier evaluated under that same distribution. The guarantee is about minimum error, not perfect prediction. Because f_D is defined using D, an unknown distribution prevents us from directly utilizing the ideal predictor.

Key Takeaways

  • Classification error measures how often a classifier's predicted binary label differs from the true label under a distribution.
  • The Bayes optimal predictor f_D is defined relative to a particular distribution D.
  • For every classifier g, f_D has error no greater than g under the same D.
  • Bayes optimality means the lowest achievable error, not necessarily zero error.
  • Unknown D makes the Bayes optimal predictor an ideal benchmark that cannot be directly utilized.