Theorem 26.15
Theorem 26.12 and Theorem 26.15 have bounds that look similar apart from an extra log(d) factor in Theorem 26.15.
The Similarity Trap
The bounds in Theorem 26.12 and Theorem 26.15 look very similar. The most visible difference is an extra log(d) factor in Theorem 26.15. That visual similarity can be misleading, because B and R do not mean the same things in the two theorems. Before applying either result, identify what each parameter constrains.
Read B and R by their roles, not only by their positions in the bound. B concerns the predictor parameter w, while R concerns the instances.
What B Constrains
In Theorem 26.12, B imposes an ℓ2 constraint on w. In Theorem 26.15, B imposes an ℓ1 constraint on w. Thus, B refers to the same predictor parameter in both theorems, but it limits that parameter using a different norm.
The ℓ1 constraint is stronger than the ℓ2 constraint. For the same stated limit B, the set of w values that satisfy the ℓ1 constraint is more restricted than the set allowed by the ℓ2 constraint. Therefore, choosing Theorem 26.15 is not merely replacing one symbol for another: it changes the assumption made about the predictor.
Tracking B across the theorems
Suppose you are deciding which theorem matches prior knowledge about a good predictor w. What does B mean in each possible choice?
Check Theorem 26.12: B describes an ℓ2 constraint placed on w.
Check Theorem 26.15: B describes an ℓ1 constraint placed on w.
Compare the assumptions: The ℓ1 constraint is stronger, so the two choices do not make the same assumption about which predictors are allowed.
B is a bound on w in both theorems, but the norm used for that bound determines the theorem's predictor assumption.
What R Says About Instances
R does not constrain w. It describes a norm assumption on the instances. In Theorem 26.12, R represents a low ℓ2-norm assumption on the instances. In Theorem 26.15, R represents a low ℓ∞-norm assumption on the instances.
The two meanings of R must remain separate from the two meanings of B. B changes the allowed predictor parameter, while R changes the assumed geometry of the input instances. The source describes the low ℓ∞-norm assumption as weaker than the low ℓ2-norm assumption.
Comparing the Two Bounds
| Feature | Theorem 26.12 | Theorem 26.15 |
|---|---|---|
| Constraint controlled by B | ℓ2 constraint on w | ℓ1 constraint on w |
| Assumption represented by R | Low ℓ2 norm on instances | Low ℓ∞ norm on instances |
| Relative strength of predictor constraint | Weaker than the ℓ1 constraint | Stronger than the ℓ2 constraint |
| Visible difference in the bound | No extra log(d) factor described here | Extra log(d) factor |
The parameters have different meanings even though the theorem bounds look similar.
The extra log(d) factor in Theorem 26.15 is the prominent visible difference between the bounds. However, comparing only that factor is incomplete. A correct comparison also asks whether the problem supports the stronger ℓ1 constraint on w and whether the instances fit the low ℓ∞-norm assumption represented by R.
Selecting a Perspective
Selection should begin with prior knowledge about the problem, not with the superficial similarity of the displayed bounds. Ask two questions: what norm assumption is credible for the instances, and what norm constraint is credible for a good predictor w?
A theorem-selection decision
You know that the likely good predictor is expected to satisfy the stronger ℓ1 constraint, and your instance knowledge supports the low ℓ∞-norm perspective. Which theorem's assumptions should you investigate first?
Inspect the predictor: Theorem 26.15 uses the ℓ1 constraint on w, which is stronger than the ℓ2 constraint used in Theorem 26.12.
Inspect the instances: Theorem 26.15 represents R as a low ℓ∞-norm assumption on the instances.
Account for the bound: Theorem 26.15 has an extra log(d) factor, so the visual form of the bound should not be compared without checking whether its assumptions fit.
Investigate Theorem 26.15 first, because both the predictor and instance assumptions match that theorem's perspective.
Common Reading Errors
Treating B as though it has the same norm meaning in both theorems.
Theorem 26.15 uses an ℓ1 constraint on w.
Fix:
Identify the norm attached to B separately in each theorem.Treating R as a second constraint on w.
R represents a norm assumption on the instances.
Fix:
Keep B associated with w and R associated with the instances.Assuming the similar-looking bounds make the two theorems interchangeable.
The meanings of both B and R change between the theorems.
Fix:
Compare the predictor constraint, the instance assumption, and the displayed factor together.Forgetting that the ℓ1 constraint is stronger than the ℓ2 constraint.
The ℓ1 constraint allows a more restricted set of predictors.
Fix:
When evaluating Theorem 26.15, explicitly check whether the stronger predictor assumption is justified.
Apply the Distinction
A theorem statement contains a parameter B and a parameter R. You are comparing Theorem 26.12 with Theorem 26.15. Write a four-part identification: the meaning of B in Theorem 26.12, the meaning of B in Theorem 26.15, the meaning of R in Theorem 26.12, and the meaning of R in Theorem 26.15. Then state which theorem uses the stronger constraint on w and which theorem contains the extra log(d) factor.
Hints
- B always concerns w, but the norm changes.
- R concerns the instances, not w.
- Check the theorem comparison for the extra factor.
Key Takeaways
- Theorem 26.12 uses B for an ℓ2 constraint on w; Theorem 26.15 uses B for an ℓ1 constraint on w.
- The ℓ1 constraint on w is stronger, so it permits a more restricted set of predictors than the ℓ2 constraint.
- R describes the instances: low ℓ2 norm in Theorem 26.12 and low ℓ∞ norm in Theorem 26.15.
- Theorem 26.15 has an extra log(d) factor, but the deeper comparison requires checking both B and R.
- Choose the theorem whose predictor and instance assumptions match the prior knowledge available for the problem.
Key Takeaways
- B has different norm meanings in the two theorems: ℓ2 in Theorem 26.12 and ℓ1 in Theorem 26.15.
- The ℓ1 constraint is stronger than the ℓ2 constraint on w.
- R refers to the instances, with a low ℓ2-norm assumption in Theorem 26.12 and a low ℓ∞-norm assumption in Theorem 26.15.
- The extra log(d) factor is only one part of the comparison; the assumptions represented by B and R also change.
- The correct theorem depends on prior knowledge about the instances and the likely good predictor.