Logistic Regression
Try it: Logistic Regression
How logistic regression turns the linear score ⟨w,x⟩ + b into a probability h_w(x) = σ(⟨w,x⟩ + b), how gradient descent on the average logistic loss log(1 + exp(−y(⟨w,x⟩ + b))) moves the boundary, and how the decision threshold decides which points count as training errors.
How it works
- The linear stage computes a real score ⟨w,x⟩ + b for every point; the sigmoid stage maps it into [0, 1] as h_w(x), read as the probability that the label is +1.
- Each labelled point (y ∈ {±1}) costs log(1 + exp(−y(⟨w,x⟩ + b))): small when y·score is large and positive, large when the score contradicts the label.
- ERM averages these losses into L_S(w); gradient descent starts at w = 0 (every h = 0.5, every loss ln 2) and repeats w ← w − η∇L_S(w).
- Or place the boundary yourself: the line through handles A and B, with steepness ‖w‖, and read off every probability and loss.
- Finally the threshold t: predict +1 when h_w(x) ≥ t, i.e. when ⟨w,x⟩ + b ≥ ln(t/(1 − t)); count the training errors.
Default run (57 steps): Start at w = 0, b = 0: every score is 0, so every h_w(x) = σ(0) = 0.5 and every logistic loss is ln 2 ≈ 0.693. Gradient descent now steps w ← w − η∇L_S(w) with η = 0.5. … Trained: w = (0.34, 0.478), b = -0.521, average logistic loss 0.51. Threshold: predict +1 when h_w(x) ≥ 0.5 ⇔ ⟨w,x⟩ + b ≥ ln(0.5/0.5) = 0. That gives 2 training errors (p_6, p_12).
Simplified: Two features on a half-unit grid, at most 14 points, full-batch gradient descent with a fixed step size and no regularisation. On separable data the weights keep growing, so training simply stops after the chosen number of iterations (or when the gradient is below 1e-9). Some intermediate iterations are grouped into one step when there are more than 60.
Loading the simulation…