Concepts / Probability Distributions in Learning

Probability Distributions in Learning

Classifier error measures the probability of an incorrect prediction on a randomly drawn instance.

  • Programming

A Random Prediction Trial

Imagine repeatedly receiving an instance selected according to an underlying probability distribution. A classifier receives that instance and produces a prediction. The prediction is then compared with the correct label. Classifier error describes how often this process produces an incorrect prediction in probability terms. It is not the name of one particular failed prediction; it is the probability of landing in the incorrect-prediction outcome.

A single wrong prediction is an event. Classifier error is the probability of that event when instances are drawn according to the underlying distribution.

From Instance to Error

drawsclassified byhassupplies h(x)supplies f(x)agreesdisagreesDistribution Ddraw probabilitiesInstance xrandomly drawnPrediction h(x)classifier outputLabel comparisonh(x) versus f(x)Correct predictionCorrect label f(x)labeling functionIncorrect prediction
What happens when an instance is randomly drawn, passed to the classifier, and compared with its correct label?

For an instance x, h(x) denotes the classifier's prediction and f(x) denotes the correct label. The trial is correct when these agree and incorrect when they disagree. To determine the classifier's overall error, consider random instances generated according to D and determine the probability of the disagreement cases.

Counting Incorrect Cases

When a small collection of instances is equally likely, classifier error can be found by counting the cases where h(x) and f(x) disagree. Divide the number of incorrect cases by the total number of cases. This is a counting version of the probability definition.

InstanceClassifier prediction h(x)Correct label f(x)Outcome
x1AACorrect
x2BBCorrect
x3ABIncorrect
x4BBCorrect

Generated illustrative collection with four equally likely instances.

Error in an Equally Likely Collection

A collection contains four equally likely instances. The classifier and correct labeling function disagree on x3 only. What is the classifier's error?

Identify disagreements: Compare h(x) with f(x) for every instance. Only x3 is an incorrect-prediction case.

Count all cases: There are four possible instances in the collection.

Convert the count to probability: One incorrect case among four equally likely cases gives an error probability of 1 out of 4, or 0.25.

The classifier's error is 0.25, or 25 percent, for this illustrative distribution.

agreesagreesdisagreesagrees1 of 4x1h=A; f=A3 correct casesError 0.25x2h=B; f=B1 incorrect casex3h=A; f=Bx4h=B; f=B
Which instances are misclassified, and how does counting them produce the classifier's error probability?

Why Distribution and Labels Matter

The probability distribution D matters because it determines how likely each instance is to be drawn. Therefore, raw disagreement counts are enough only in a setting where the instances have equal likelihood. With a non-uniform distribution, the error is determined by the probability assigned to the incorrect-prediction cases, not by their raw count alone.

Generated example: suppose two instances are incorrect cases, but one is much more likely to be drawn than the other. Their contribution to classifier error is not determined merely by the fact that there are two of them. The distribution's probabilities determine how much those two cases contribute to the overall error.

The correct labeling function matters as well. The classifier is judged by comparing h(x) with f(x). If the correct labels change while the classifier's predictions stay the same, the set of disagreements can change. If the distribution changes while the labels and predictions stay the same, the likelihood of drawing those disagreement cases can change. Classifier error therefore depends on both which instances are likely to be drawn and which labels count as correct.

Three Names for One Quantity

TermMeaning in this framework
Generalization errorThe distribution-based probability that the classifier makes an incorrect prediction.
RiskThe same distribution-based classifier error quantity.
True errorThe same distribution-based classifier error quantity.

These terms do not describe three different calculations here. Each refers to the probability, under the relevant distribution, that h(x) disagrees with f(x).

Common Counting Mistakes

  • Treating one failed prediction as the classifier's entire error.

    That is one incorrect-prediction outcome, not the probability of reaching that outcome over randomly drawn instances.

    Fix: Consider the distribution of possible instances and determine the probability of all cases where h(x) and f(x) disagree.

  • Counting disagreements without considering their probabilities.

    With a non-uniform distribution, the raw number of incorrect cases does not by itself determine the error.

    Fix: Use the probability assigned by D to the incorrect-prediction cases.

  • Comparing the prediction with the wrong reference.

    Error is defined by the comparison between h(x) and f(x).

    Fix: For each instance, identify both the classifier prediction and the correct label before deciding whether the case is incorrect.

  • Treating generalization error, risk, and true error as different quantities in this framework.

    The source defines these as names for the same distribution-based error.

    Fix: Recognize that each term refers to the probability of an incorrect prediction under the distribution.

Check Your Understanding

What do you think happens?

A collection has five equally likely instances. The classifier disagrees with the correct labeling function on two of them. What is the classifier's error for this collection?

  • 0.2
  • 0.4
  • 0.5
  • 2
Reveal answer

Answer: 0.4

There are two incorrect-prediction cases among five equally likely cases, so the probability of an incorrect prediction is 2 out of 5, or 0.4.

MEDIUM

Explain what could happen to classifier error if the classifier's predictions and correct labels stayed fixed, but the distribution changed so that an incorrect-prediction instance became more likely to be drawn.

Hints
  • First identify whether the set of disagreements changes.
  • Then ask how the probability assigned to the disagreement cases changes.

Key Takeaways

  1. Classifier error is the probability of an incorrect prediction on a randomly drawn instance.
  2. For an instance x, the relevant comparison is between h(x), the prediction, and f(x), the correct label.
  3. The distribution D determines how likely each instance is to be drawn, so non-uniform probabilities matter.
  4. The correct labeling function matters because it determines which predictions count as correct or incorrect.
  5. Generalization error, risk, and true error are names for the same distribution-based quantity in this framework.

Key Takeaways

  • Classifier error measures the probability of disagreement between h(x) and f(x) for an instance drawn according to D.
  • In an equally likely collection, error can be calculated as the fraction of instances on which the classifier is incorrect.
  • With unequal probabilities, sum the probability assigned to the incorrect-prediction cases rather than relying on their count alone.
  • Generalization error, risk, and true error refer to the same distribution-based error quantity in this framework.