Playground / Gradient Descent: Full-Batch, Mini-Batch and Stochastic

Race full-batch GD against SGD

Gradient Descent: Full-Batch, Mini-Batch and Stochastic

Interactive lab

Try it: Gradient Descent: Full-Batch, Mini-Batch and Stochastic

How gradient descent, mini-batch gradient descent and stochastic gradient descent differ only in how many examples estimate the gradient — and how that trades noisy steps for many cheap updates.

How it works

  1. The loss is the mean squared error of the line ŷ = w·x + b over 12 points.
  2. At the start of every epoch, shuffle the examples (seeded) and cut them into batches.
  3. For each batch, average the gradient 2(ŵx + b − y)·(x, 1) over its examples.
  4. Update (w, b) ← (w, b) − learning rate × gradient; the path is drawn over the loss contours.
  5. Batch size 12 is full-batch gradient descent, 1 is SGD; too large a learning rate overshoots and diverges.

Default run (32 steps): mini-batch gradient descent (batch size 4): start at w = -1, b = 3 (MSE 9.823), learning rate 0.1, 3 updates per epoch. The least-squares optimum is w* = 1.595, b* = 0.404. … After 10 epochs (30 updates): w = 1.608, b = 0.468, MSE 0.5665 vs the optimum 0.5613 (distance 0.066). Averaged iterate w̄ = 1.106, b̄ = 0.996.

Simplified: A two-parameter linear-regression loss on 12 fixed points with a constant learning rate; real SGD runs on large datasets and often decays the learning rate.

Educational simulation

Loading the simulation…