Hypothesis
Classifier error measures the probability of an incorrect prediction on a random data point.
A Prediction Is Tested on Random Data
A classifier is not evaluated by looking at only one prediction. Instead, imagine selecting a data point according to an underlying distribution D. The correct labeling function f supplies the correct label for that point, while a hypothesis or classifier h supplies its prediction. The classifier error is the probability that h gives a different label from f.
The Three Roles in Error
The distribution D determines which data points can be selected and how likely they are. The function f represents the correct labeling rule. The hypothesis h is the classifier being evaluated. An error occurs precisely when h(x) does not equal f(x). Therefore, classifier error depends on both the distribution D and the correct labeling function f, as well as on the behavior of h on the possible points.
Classifier error is the probability, over a randomly selected data point from D, that the classifier prediction h(x) differs from the correct label f(x).
Generalization error, risk, and true error are synonymous names for this same quantity in this context: the probability of an incorrect prediction under the data distribution.
Adding Probabilities of Mistakes
To calculate error for a small, discrete set of possible data points, inspect each point and find where h and f disagree. Then add the probabilities assigned by D to those incorrect points. The result is the probability that a randomly selected point is misclassified.
A Four-Point Error Calculation
Suppose D selects four possible points with probabilities 0.10, 0.20, 0.30, and 0.40. For the first and third points, h agrees with f. For the second and fourth points, h disagrees with f. What is the classifier error?
Find the incorrect points: The incorrect predictions are the second and fourth points because h(x) does not equal f(x) on those points.
Read their probabilities: The distribution assigns probability 0.20 to the second point and 0.40 to the fourth point.
Add the error probabilities: The probability of selecting an incorrect point is 0.20 plus 0.40, which equals 0.60.
The classifier error is 0.60, or 60 percent.
ERM Searches a Finite Class
A learner may be restricted to a fixed collection of predictors called a hypothesis class. When this class is finite, it contains a finite number of candidate hypotheses, even if the collection is still large. Empirical risk minimization, or ERM, examines how every candidate performs on the training sample and chooses a hypothesis with minimum empirical risk.
Restricting the learner to a finite class gives it a controlled candidate set. The class size matters because a larger collection gives the learner more hypotheses to compare and more opportunities to find a candidate that fits the observed sample unusually well. The amount of training data required therefore grows in relation to the size of the hypothesis class.
Realizability Creates a Perfect Candidate
The realizability assumption states that the hypothesis class contains a hypothesis h* with zero true risk. In other words, at least one available hypothesis labels every possible point correctly according to the correct labeling function.
Because h* has zero true risk, a random sample drawn from D and labeled by f has empirical risk zero for h* with probability 1. Thus ERM has at least one zero-error candidate on the observed sample. Since ERM chooses a hypothesis with minimum empirical risk, every ERM result hS also has empirical risk zero under realizability.
Training Error Versus True Error
Empirical risk describes how a hypothesis performs on the particular training sample S. True risk, also called generalization error or true error, describes how it performs on randomly drawn data points from D. A hypothesis can make no mistakes on S while still disagreeing with f on points that were not represented in S. Therefore, zero empirical risk is not the same statement as low true risk.
Generated example: imagine two hypotheses that both label every point in a small training sample correctly. One of them also behaves correctly on points drawn from D, while the other disagrees with f on some points that were absent from the sample. Both have zero empirical risk, but their true risks can differ. The training result alone cannot identify the difference.
| Quantity | What is evaluated | What it tells you |
|---|---|---|
| Empirical risk | A hypothesis on the observed training sample S | How well the hypothesis fits the sample |
| True risk | A hypothesis on randomly drawn points from D | The probability of an incorrect prediction |
| Generalization error | A hypothesis under the data distribution | Another name for true risk in this context |
Class Size and Sample Requirements
For a finite hypothesis class, the learner compares a fixed number of candidates. The class size is central to the question of whether a hypothesis selected using the training sample will also perform well beyond that sample. As the number of available hypotheses grows, the learner has a larger candidate set and more opportunity to find a hypothesis that fits the observed sample. The required amount of training data grows in relation to the size of the class.
When reasoning about an ERM result, ask two separate questions: how well did the selected hypothesis perform on the observed sample, and how much evidence is available for judging performance under D? The first concerns empirical risk. The second is connected to class size and the amount of training data.
Mistakes in Risk Reasoning
Treating one incorrect prediction as the classifier error.
Classifier error is a probability over randomly selected data points from D, not the result of one observation.
Fix:
Consider all possible points and add the probabilities of the points where h(x) differs from f(x).Ignoring D or f when discussing error.
The measurement depends on both D and f.
Fix:
State which points D can select and which labels f assigns before evaluating h.Equating zero empirical risk with zero true risk.
The hypothesis may disagree with f on points not represented in S.
Fix:
Separate the sample-based result from the true risk under D.Assuming realizability means every hypothesis is perfect.
Realizability places at least one zero-true-risk hypothesis h* inside H; it does not make every candidate perfect.
Fix:
Use h* as the guaranteed perfect candidate and remember that ERM chooses among all candidates by empirical risk.Assuming a finite hypothesis class automatically removes all overfitting.
Finiteness provides a controlled candidate set, but the relationship between class size and training-data amount still matters.
Fix:
Consider both the size of H and the amount of training data.
Check Your Understanding
A distribution D selects three possible points with probabilities 0.25, 0.35, and 0.40. A hypothesis h agrees with the correct labeling function f on the first and third points but disagrees on the second. What is the classifier error? Then explain whether a different hypothesis with zero empirical risk on a training sample must have zero true risk.
Hints
- Identify the points where h(x) does not equal f(x).
- Add the probabilities of only those incorrect points.
- For the second part, distinguish the observed sample from new points drawn according to D.
What do you think happens?
Under realizability, ERM has access to h*, a hypothesis with zero true risk. What empirical-risk conclusion follows for an ERM result hS?
Reveal answer
Answer: hS must have empirical risk zero
Under realizability, h* has empirical risk zero on the random sample with probability 1. ERM chooses a hypothesis with minimum empirical risk, so its result also has empirical risk zero. This does not by itself establish zero true risk for hS.
The Complete Reasoning Chain
- D selects a random data point, and f supplies its correct label.
- A classifier h makes an error exactly when h(x) differs from f(x).
- The classifier error is the probability of selecting a point where those labels differ.
- Generalization error, risk, and true error name this same distribution-level quantity in this setting.
- ERM evaluates each candidate in a finite hypothesis class on the training sample and selects one with minimum empirical risk.
- Under realizability, a zero-true-risk hypothesis h* belongs to the class, so h* has zero empirical risk on the random sample with probability 1 and ERM also achieves zero empirical risk.
- Zero empirical risk is still a statement about the sample; true risk concerns performance on points drawn from D.
- The larger the hypothesis class, the more the required training-data amount matters for controlling the gap between sample performance and performance under D.
Key Takeaways
- Classifier error is the probability that h(x) differs from the correct label f(x) on a point selected from D.
- The distribution D determines how likely each possible point is, while f determines the correct label.
- Generalization error, risk, and true error refer to the same quantity here.
- ERM selects a minimum-empirical-risk hypothesis from a finite hypothesis class.
- Realizability guarantees zero empirical risk for ERM with probability 1, but zero empirical risk is not automatically low true risk; class size and training-data amount remain important.