Playground / Backpropagation, Step by Step

Follow the gradient back through a tiny network

Backpropagation, Step by Step

Interactive lab

Try it: Backpropagation, Step by Step

A forward pass through a 2-2-1 network, then the backward pass that applies the chain rule node by node to get dL/dw for every weight and bias, and one gradient-descent update that changes the loss.

How it works

  1. Forward: compute each hidden pre-activation and activation, then the sigmoid output y and the loss.
  2. Backward from the loss: dL/dz_o = 2(y - t) y(1 - y) for squared error, or y - t for binary cross-entropy.
  3. Output weights: dL/dv_j = dL/dz_o * h_j; the error sent back to h_j is dL/dz_o * v_j.
  4. Hidden units: multiply by the activation's local derivative, then dL/dw_ji = dL/dz_j * x_i.
  5. Update every parameter p <- p - lr * dL/dp and run the forward pass again to see the new loss.

Default run (12 steps): Inputs x = (1, 0.5), target t = 1. Hidden activation sigmoid, output sigmoid, loss: squared error L = (y - t)^2. First the forward pass. … New forward pass with the updated weights: y = 0.6086, loss 0.203 -> 0.1532 (lower: the step moved downhill).

Simplified: One training example, two inputs, two hidden units, one output unit. Binary cross-entropy is evaluated from the output logit (same value, no log(0)). Toy dimensions, not a full real model.

Educational simulation

Loading the simulation…