Concepts / Theorem 26.12

Theorem 26.12

Theorem 26.12 and Theorem 26.15 have bounds that look similar apart from an extra log(d) factor in Theorem 26.15.

  • Programming

The Similarity Trap

Theorem 26.12 and Theorem 26.15 have bounds that look very similar. The main visible difference is that Theorem 26.15 contains an extra log(d) factor. That visual similarity is not enough to decide which theorem applies, because the symbols B and R describe different assumptions in the two results. The reliable method is to read both parameters by meaning: first identify the constraint on the predictor parameter w, then identify the norm assumption on the instances.

uses assumptionsuses assumptionsTheorem 26.12B: ℓ2 constraint on w; R:low ℓ2 normSimilar boundNo extra log(d) factordescribedTheorem 26.15B: ℓ1 constraint on w; R:low ℓ∞ normSimilar boundExtra log(d) factor
What changes between the two bounds, and why can their visual similarity be misleading?

Tracking B and R

In Theorem 26.12, B imposes an ℓ2 constraint on w. In Theorem 26.15, B imposes an ℓ1 constraint on w. Thus, B refers to a constraint on the predictor parameter in both theorems, but the norm used for that constraint changes. The parameter R is different in another way: it describes a norm assumption on the instances, not on w. In Theorem 26.12, R represents a low ℓ2-norm assumption on the instances. In Theorem 26.15, R represents a low ℓ∞-norm assumption on the instances.

TheoremMeaning of BMeaning of RVisible bound difference
Theorem 26.12ℓ2 constraint on wLow ℓ2-norm assumption on instancesOne of the similar-looking bounds
Theorem 26.15ℓ1 constraint on wLow ℓ∞-norm assumption on instancesExtra log(d) factor

The same letters do not represent the same norm assumptions across the two theorems.

Why the Predictor Constraint Matters

The ℓ1 constraint on w is stronger than the ℓ2 constraint on w. In practical terms, the ℓ1 perspective allows a more restricted collection of predictor parameters than the corresponding ℓ2 perspective. Therefore, Theorem 26.15 is not simply Theorem 26.12 with an extra log(d) factor added to an otherwise identical setting. Its assumption on w is stronger, because it uses an ℓ1 constraint rather than an ℓ2 constraint.

permitspermitsℓ2 constraintConstraint on wAllowed wLess restricted perspectiveℓ1 constraintConstraint on wAllowed wMore restricted perspective
How do the allowed predictor parameters differ when the constraint on w changes from ℓ2 to ℓ1?

Reading B Without Confusing the Theorems

Suppose you are deciding whether the predictor-side assumption matches Theorem 26.12 or Theorem 26.15. What should you inspect first?

Identify the object controlled by B: B controls the predictor parameter w in both theorems.

Identify the norm: Theorem 26.12 uses an ℓ2 constraint on w, while Theorem 26.15 uses an ℓ1 constraint on w.

Compare the strength: The ℓ1 constraint is stronger than the ℓ2 constraint, so the two theorems make different predictor-side assumptions.

Avoid judging from the displayed bound alone: The bounds look similar, but the meaning of B changes with the theorem.

A correct comparison begins with the norm imposed on w, not with the superficial similarity of the two bounds.

Reading R on the Instance Side

R must be read on the instance side of the result. For Theorem 26.12, a small R represents a low ℓ2-norm assumption on the instances. For Theorem 26.15, a small R represents a low ℓ∞-norm assumption on the instances. The source describes the low ℓ∞-norm assumption as weaker than the low ℓ2-norm assumption. Consequently, two theorems can have similar-looking bounds while relying on different assumptions about the data they receive.

summarizessummarizesR in Theorem 26.12Low ℓ2 norm of instancesInstance norm profileℓ2-basedR in Theorem 26.15Low ℓ∞ norm of instancesInstance norm profileℓ∞-based
What does R measure under each theorem, and why can equal-looking R values represent different instance assumptions?

Selecting the Constraint Perspective

The theorem should be selected from prior knowledge about the problem, not from the appearance of the final bound. Ask two questions. First, what kind of predictor parameter w is expected to describe a good predictor: one that fits the ℓ2-constrained perspective, or one that fits the stronger ℓ1-constrained perspective? Second, what is known about the instances: is a low ℓ2-norm assumption or a low ℓ∞-norm assumption more appropriate? Theorem 26.12 pairs an ℓ2 constraint on w with a low ℓ2-norm assumption on instances. Theorem 26.15 pairs an ℓ1 constraint on w with a low ℓ∞-norm assumption on instances.

inspectℓ2 perspectiveℓ1 perspectiveif instance assumption fitsif instance assumption fitsKnown problemstructurePrior knowledgeExpected goodpredictorWhich constraint on w fits?ℓ2 on wCheck low ℓ2 norm oninstancesTheorem 26.12B: ℓ2; R: low ℓ2ℓ1 on wCheck low ℓ∞ norm oninstancesTheorem 26.15B: ℓ1; R: low ℓ∞
How should prior knowledge about the predictor and the instances guide the theorem choice?

Before applying either theorem, write a two-line assumption check: B describes the norm constraint on w; R describes the norm assumption on the instances. Then verify whether the pair is ℓ2 with low ℓ2, or ℓ1 with low ℓ∞. This prevents the extra log(d) factor from becoming the only feature you compare.

Common Reading Errors

  • Assuming B has the same meaning in both theorems.

    Theorem 26.15 uses an ℓ1 constraint on w.

    Fix: Identify the norm attached to w separately for each theorem.

  • Treating R as a constraint on w.

    R describes a norm assumption on the instances.

    Fix: Assign B to w and R to the instances before comparing the results.

  • Choosing a theorem only because its displayed bound looks better.

    The two bounds look similar, but their B and R parameters do not mean the same thing.

    Fix: Check both the predictor constraint and the instance norm assumption.

  • Assuming that an ℓ1 constraint is interchangeable with an ℓ2 constraint.

    The ℓ1 constraint restricts w more strongly than the ℓ2 constraint.

    Fix: Use prior knowledge about the expected predictor before selecting the constraint perspective.

Assumption Check

MEDIUM

You are given a learning problem in which prior knowledge suggests that the expected good predictor should be evaluated through the stronger ℓ1 constraint on w. The instances are also expected to satisfy the low ℓ∞-norm assumption. Which theorem perspective should you investigate, and what do B and R mean in that perspective?

Hints
  • Find the theorem whose B parameter uses an ℓ1 constraint on w.
  • Then identify the instance norm represented by R in that same theorem.

Practice Solution

Choose the theorem perspective for a good predictor described by an ℓ1 constraint on w and instances described by a low ℓ∞ norm.

Match the predictor constraint: The ℓ1 constraint on w identifies Theorem 26.15.

Match the instance assumption: Theorem 26.15 represents R as a low ℓ∞-norm assumption on the instances.

Interpret the visible difference: Theorem 26.15 has the extra log(d) factor, but that factor must be considered together with its different B and R assumptions.

Investigate Theorem 26.15: B imposes an ℓ1 constraint on w, and R represents a low ℓ∞-norm assumption on the instances.

Key Takeaways

  • Theorem 26.12 uses an ℓ2 constraint on w, while Theorem 26.15 uses an ℓ1 constraint on w.
  • The ℓ1 constraint on w is stronger than the ℓ2 constraint.
  • R describes the instances, not the predictor parameter: it represents low ℓ2 norm in Theorem 26.12 and low ℓ∞ norm in Theorem 26.15.
  • The bounds look similar, except for the extra log(d) factor in Theorem 26.15, but their assumptions differ.
  • Choose between the theorem perspectives by matching prior knowledge about the good predictor and the instances to the corresponding B and R meanings.