Loss Functions of the Margin
Try it: Loss Functions of the Margin
How the 0-1, hinge, logistic and squared losses each score a classifier's margin z = y(⟨w,x⟩ + b), and why hinge (and base-2 logistic) are convex surrogates that upper-bound the 0-1 loss.
How it works
- For each point compute ⟨w,x⟩ + b and the margin z = y(⟨w,x⟩ + b): positive means correctly classified.
- 0-1 loss: 1 if z ≤ 0, else 0.
- Hinge loss: max(0, 1 − z); logistic loss: log(1 + e^(−z)); squared loss: (1 − z)², which equals (y − (⟨w,x⟩ + b))² because y² = 1.
- Average each loss over the points to get the empirical risks.
- Hinge, squared and base-2 logistic are ≥ the 0-1 loss everywhere; natural-log logistic dips below it near z = 0.
Default run (9 steps): Classifier w = (0.5, 0.866), b = 0. For each point the margin z = y(⟨w,x⟩ + b) is positive when it is on its correct side; every loss is a function of z. … Averages (empirical risks): 0-1 0.286, hinge 0.338, logistic 0.333, squared 1.887. Hinge and squared loss are ≥ the 0-1 loss at every point; natural-log logistic is below it at 2 points (ln 2 ≈ 0.69 < 1 at z = 0) — use base 2 for a true upper bound.
Simplified: Linear classifier in 2-D with at most 10 points; the boundary case z = 0 counts as a mistake.
Loading the simulation…