Backpropagation, Step by Step
Interactive lab
Try it: Backpropagation, Step by Step
A forward pass through a 2-2-1 network, then the backward pass that applies the chain rule node by node to get dL/dw for every weight and bias, and one gradient-descent update that changes the loss.
How it works
- Forward: compute each hidden pre-activation and activation, then the sigmoid output y and the loss.
- Backward from the loss: dL/dz_o = 2(y - t) y(1 - y) for squared error, or y - t for binary cross-entropy.
- Output weights: dL/dv_j = dL/dz_o * h_j; the error sent back to h_j is dL/dz_o * v_j.
- Hidden units: multiply by the activation's local derivative, then dL/dw_ji = dL/dz_j * x_i.
- Update every parameter p <- p - lr * dL/dp and run the forward pass again to see the new loss.
Default run (12 steps): Inputs x = (1, 0.5), target t = 1. Hidden activation sigmoid, output sigmoid, loss: squared error L = (y - t)^2. First the forward pass. … New forward pass with the updated weights: y = 0.6086, loss 0.203 -> 0.1532 (lower: the step moved downhill).
Simplified: One training example, two inputs, two hidden units, one output unit. Binary cross-entropy is evaluated from the output logit (same value, no log(0)). Toy dimensions, not a full real model.
Educational simulation
Loading the simulation…