Gradient Descent
Interactive lab
Try it: Gradient Descent
How gradient descent minimises a loss by repeatedly stepping against the gradient, and how the learning rate controls whether it converges, oscillates or diverges.
How it works
- The loss is f(x) = (x − 3)² + 1, lowest at x = 3.
- At the current x, compute the gradient f′(x) = 2(x − 3).
- Update x ← x − learning rate × gradient.
- Repeat. Small learning rates creep in, good ones converge quickly, rates above 1 overshoot further each step and diverge.
Default run (40 steps): Start at x = -3. Loss f(x) = (x − 3)² + 1 = 37. Learning rate 0.1. … Update 39: gradient f′ = 2(x − 3) = -0.0025; step = −0.1 × -0.0025 = 0.0002; new x = 2.999; loss 1 → 1. x is within 0.001 of the minimum at x = 3 — converged.
Simplified: One parameter and a perfectly smooth loss. Real models optimise millions of parameters on noisy losses.
Educational simulation
Loading the simulation…