Jacobian Matrix
Backpropagation calculates the gradient of a loss with respect to network weights for an example.
Tracing Responsibility Through a Network
Backpropagation finds how the loss for one example depends on the network's weights. The difficult part is that an early weight does not affect the loss directly. It changes one layer's result, which changes later layers, which eventually changes the loss. Backpropagation follows that dependency chain in two directions: forward to compute the quantities that actually occurred, and backward to determine how much each earlier quantity contributed to the loss.
Forward Quantities First
The forward pass evaluates the network from the input toward the output. The input is treated as the initial output, o₀ = x. At layer t, the network forms a pre-activation from the previous output and the layer's weights, written in the source as aₜ = W₍ₜ₋₁₎o₍ₜ₋₁₎. It then applies the activation function to obtain oₜ = σ(aₜ). This pass produces the pre-activations and outputs that the backward calculation will later need.
Following one layer forward
A layer receives the output from the preceding layer and applies its weights followed by an activation function. What quantities must be available before backpropagation can use this layer?
Receive the source output: The layer begins with o₍ₜ₋₁₎, the output produced by the preceding layer.
Form the pre-activation: The weights combine with the source output to produce aₜ = W₍ₜ₋₁₎o₍ₜ₋₁₎.
Apply the activation: The activation function transforms aₜ into the layer output oₜ = σ(aₜ).
Retain the forward values: The backward pass needs the source output and the destination pre-activation when it forms local derivatives for the weights.
The forward pass must provide the intermediate pre-activations and outputs, not only the final prediction.
Jacobian Entries as Local Maps
A Jacobian organizes partial derivatives by output component and input variable. Each entry describes how one output component changes with respect to one input variable. Its shape records how many outputs depend on how many inputs.
For a linear function f(w) = Aw, the Jacobian is A itself. This is a useful reference case: the matrix already records how every output component responds to every input component. For an element-wise sigmoid transformation, the Jacobian is diagonal. Its diagonal entries contain the individual sigmoid derivatives, while its off-diagonal entries are zero because each output component uses its corresponding input component locally.
Multiplying Layerwise Effects
A neural network is a composition of functions. One layer transforms its input, the next layer transforms that result, and so on. The chain rule connects these transformations: the derivative of the full composition is obtained by multiplying the Jacobians of the component functions in the appropriate order. Backpropagation is the organized use of this rule from the final layer toward the first.
Why an early weight needs later information
An early weight changes the output of its layer. That output then passes through another layer before affecting the loss. What kinds of factors must a backpropagation calculation combine?
Local effect at the early edge: The calculation considers how changing the early weight changes the destination neuron's pre-activation and output.
Effect of the destination: The activation derivative at that destination determines how a pre-activation change becomes an output change.
Downstream responsibility: The error information propagated from later layers records how strongly the destination output affects the loss.
Chain the effects: The local and downstream effects are combined through the chain rule rather than recomputed as an unrelated full derivative for each weight.
An early weight receives responsibility from both its local effect and the later network computation.
From Error Vectors to Edge Gradients
The backward pass starts at the final layer. For the squared loss described in the source, the final error vector is δₜ = oₜ − y. Earlier error vectors combine the next layer's weights, the next layer's error vector, and the derivative of the next layer's activation. After these error vectors are available, the derivative for an edge uses three local ingredients: the error at the destination, the activation derivative at that destination, and the output at the edge's source neuron.
Finding the Faulty Checkpoint
A gradient discrepancy is easier to diagnose when the calculation is divided into checkpoints. First inspect the forward pre-activations and outputs. If they are already wrong, the problem is in the forward evaluation. If the forward values are correct, inspect the final error and the recursively computed earlier error vectors. A mismatch there points to the backward pass. If the error vectors and activation derivatives are correct but the weight gradient is still wrong, inspect how the edge derivative was assembled from the destination error, destination activation derivative, and source output.
Checking only the final gradient
Backpropagation exposes intermediate checkpoints: forward pre-activations, forward outputs, the final error, earlier error vectors, and edge derivatives.
Fix:
Compare the expected and computed values at each checkpoint, beginning with the forward pass.Ignoring stored forward quantities
The local edge calculation needs the source output and the destination pre-activation or its activation derivative.
Fix:
Retain and inspect the intermediate forward values used by the backward calculation.Treating an edge gradient as depending on only one factor
The edge derivative is formed from the destination error, the activation derivative at the destination, and the source output.
Fix:
Check all three local ingredients before concluding that the recursive backward pass is wrong.
Practice: Locate the Discrepancy
A network produces the expected forward pre-activations and outputs. Its final error vector is also correct, but an earlier error vector differs from the expected value. Which part should you inspect first, and why?
Hints
- Separate the forward checkpoints from the recursively computed error checkpoints.
- Earlier error vectors combine later-layer weights, later error information, and an activation derivative.
What do you think happens?
The forward values and final error are correct, but an earlier error vector is incorrect. Is the most likely issue in the forward pass, the backward pass, or the final edge-gradient calculation?
Reveal answer
Answer: Backward pass
The earlier error vector is produced recursively from later-layer weights, later error information, and the next activation derivative. If that vector is wrong before the edge calculation, the discrepancy is in the backward computation.
Key Takeaways
- The forward pass computes the pre-activations and outputs that describe what the network did for an example.
- The backward pass starts at the final error and propagates responsibility toward earlier layers.
- A Jacobian organizes partial derivatives between output components and input variables.
- The chain rule combines Jacobians from successive layers to obtain derivatives through the whole network.
- To diagnose a gradient discrepancy, inspect forward values, error vectors, and final edge-gradient factors as separate checkpoints.
Key Takeaways
- Backpropagation needs a forward pass to compute intermediate network quantities and a backward pass to propagate loss responsibility.
- The Jacobian is a structured map of partial derivatives from input coordinates to output coordinates.
- The chain rule multiplies local derivative information across successive layers.
- A single edge gradient combines upstream error information with the destination activation derivative and the source output.
- Gradient debugging should proceed through forward values, error vectors, and edge-gradient assembly.