Generalization Risk
i.i.d. sampling means independent examples drawn from the same distribution D.
From Sample to Risk
An ERM learner receives a training set rather than direct access to the underlying distribution D. It can examine what happened on the sampled examples, but the guarantee we want concerns the risk of the learned hypothesis under D. Generalization risk asks how likely the sampled training set is to produce a hypothesis whose risk is unacceptably large.
The central question is not whether every learned hypothesis is perfect. It is whether the probability of producing a hypothesis with risk beyond the permitted level can be bounded.
How the Training Set Is Sampled
The notation S ∼ D^m describes a random training set S containing m examples. The i.i.d. assumption means two things: the examples are independent of one another, and every example is drawn from the same distribution D. Thus, the training set is formed by repeatedly sampling from D under the same distributional source, with the samples treated as independent.
The Risk Threshold ϵ
The parameter ϵ defines how much risk is tolerated before the learner is considered to have failed. It is a boundary for the guarantee being studied. A guarantee aims to control the probability that sampling produces an output whose risk crosses this boundary, rather than claiming that every output has zero risk.
Generated example: Suppose a learning analysis chooses ϵ as its permitted risk level. A learned hypothesis with risk at most ϵ is within the guarantee's accuracy target. A learned hypothesis whose risk exceeds ϵ is outside that target and is counted as a failure for this analysis.
Correctness versus Failure
Let hS denote the hypothesis produced from the sampled training set S, and consider its risk written as L(D,f)(hS). The learner is approximately correct when L(D,f)(hS) is at most ϵ. The learner has failed with respect to this guarantee when L(D,f)(hS) exceeds ϵ. This distinction evaluates the learned hypothesis under D, not merely on the particular sample that produced it.
Classifying a Learned Hypothesis
Generated example: A training set S produces hS. In one possible outcome, L(D,f)(hS) is at most ϵ. In another possible outcome, L(D,f)(hS) is greater than ϵ. Classify both outcomes.
First outcome: Because the risk does not exceed ϵ, the hypothesis is approximately correct under the guarantee being studied.
Second outcome: Because the risk crosses the ϵ boundary, this outcome belongs to the learner-failure event.
Interpretation: The analysis is concerned with how likely the second kind of outcome is when S is sampled.
The threshold ϵ separates approximate correctness from learner failure: risk at most ϵ is acceptable, while risk greater than ϵ is unacceptable.
Combining Bad Events
To analyze failure, define the set HB of bad hypotheses: hypotheses whose risk is beyond the permitted accuracy level. The analysis can associate separate events with these bad hypotheses and then combine those events. The Union Bound says that the probability of at least one event occurring is no greater than the sum of the separate event-probability bounds. This gives an overall upper bound on failure probability.
Adding Separate Failure Bounds
Generated example: Suppose a failure analysis identifies bad events A and B. The separate probability bounds are 0.1 for A and 0.05 for B. Use the Union Bound to obtain an upper bound for the probability that at least one event occurs.
Identify the combined event: The combined failure event is A or B: at least one of the bad events occurs.
Add the separate bounds: The Union Bound combines the separate bounds by addition: 0.1 + 0.05.
Interpret the result: The probability that at least one event occurs is bounded above by 0.15.
The Union Bound gives an overall upper bound of 0.15. It is an upper bound, not necessarily the exact probability, because A and B may overlap.
Common Reasoning Errors
Treating S ∼ D^m as if the learner directly receives D.
S ∼ D^m describes a random training set of m examples drawn from D; the learner receives S.
Fix:
Separate the distribution D from the particular sample S drawn from it.Interpreting ϵ as a claim that the hypothesis must be perfect.
ϵ defines the tolerated risk boundary. The guarantee concerns whether risk crosses that boundary.
Fix:
Ask whether L(D,f)(hS) is at most ϵ, not whether it is zero.Confusing a bad training-set outcome with every possible learner output.
The analysis bounds the probability that sampling produces an unacceptable output.
Fix:
Describe failure as an event whose probability is being bounded.Assuming Union Bound addition requires disjoint events.
The Union Bound applies even when events overlap; overlap is why the sum may be larger than the exact union probability.
Fix:
Use the sum as an upper bound for the probability that at least one event occurs.
Check Your Understanding
Generated practice: A training set S is sampled according to S ∼ D^m. The resulting hypothesis hS has risk L(D,f)(hS) greater than ϵ. Is this approximate correctness or learner failure? Then suppose the failure analysis identifies three bad events with probability bounds 0.02, 0.03, and 0.01. What upper bound does the Union Bound give for the probability that at least one occurs?
Hints
- Compare L(D,f)(hS) with the threshold ϵ.
- For the combined event, add the three separate probability bounds.
What do you think happens?
What should the Union Bound provide when three bad-event probability bounds are 0.02, 0.03, and 0.01?
Reveal answer
Answer: An upper bound of 0.06
The separate bounds are added. The result is an upper bound for the probability that at least one bad event occurs, not necessarily the exact probability.
Summary
- S ∼ D^m means that the training set contains m independent examples drawn from the same distribution D.
- The parameter ϵ is the tolerated-risk threshold used to define the guarantee.
- L(D,f)(hS) at most ϵ represents approximate correctness, while L(D,f)(hS) greater than ϵ represents learner failure for the guarantee.
- Bad hypotheses can be grouped into failure events, and the Union Bound combines their separate probability bounds into an upper bound for the probability that at least one occurs.
- The resulting sum is an upper bound even when the bad events overlap.
Key Takeaways
- The i.i.d. assumption S ∼ D^m describes m independent training examples drawn from the same distribution D.
- The accuracy parameter ϵ separates acceptable risk from unacceptable risk.
- A learned hypothesis hS is approximately correct when L(D,f)(hS) is at most ϵ; exceeding ϵ is learner failure for the stated guarantee.
- The Union Bound adds separate probability bounds to control the probability that at least one bad event occurs.
- Because bad events may overlap, the Union Bound produces an upper bound rather than necessarily the exact probability.