Concepts / Chain Rule

Chain Rule

Backpropagation computes gradients that help minimize the error between predicted and actual outputs.

  • Programming

Why the Passes Matter

Backpropagation trains a neural network by helping reduce the error between its predicted output and the actual output. Its central task is to compute gradients of the loss with respect to the model parameters. To organize that work, backpropagation separates computation into two passes: a forward pass that produces the network output and a backward pass that computes the gradients needed to update the parameters.

The forward pass answers: What output does the network produce? The backward pass answers: How should error information move through the network so parameter gradients can be computed?

Forward Pass Through Layers

The forward pass begins with an input and processes it through the network's layers. At each stage, the network computes weighted sums and then activations. The resulting activated outputs become the information passed to the next layer. This continues until the network computes its final output, which is the prediction used to compare against the actual output.

enters layeractivatespasses forwardactivatesproduces predictionInputWeighted sumWeighted sumFinal outputActivationActivation
How does an input move through each layer and become the final prediction during the forward pass?

Tracing One Forward Pass

Trace the conceptual path of an input through a network with two processing stages.

Start with the input: The input enters the first layer.

Compute the first weighted sum: The first layer combines the incoming information through its weighted-sum computation.

Apply the first activation: The first layer produces an activated output that is passed to the next layer.

Repeat the layer computation: The next layer computes another weighted sum and activation using the information passed forward.

Produce the final output: The final activated result becomes the network's prediction.

The forward pass must finish computing the final output before the backward pass can begin.

Backward Error Flow

After the final output is available, the network can determine the error between the prediction and the actual output. The backward pass then moves error information from later layers toward earlier layers. The chain rule provides the structure for relating the effect at one stage to the effects of the stages that came before it.

For each step of the backward computation, the relevant information comes from the later layer's error quantity, the connecting weights, and the activation derivative. This lets the procedure propagate error information toward earlier layers and compute gradients for the connections, also called edges, in the network.

compare with actual outputbegin backward passuse local layer informationcombine through chain rulemove toward earlier layerFinal outputFinal errorLater-layer gradientEdge gradientEarlier-layergradientActivation derivative
How does the error signal move backward through each layer, and how does the chain rule support each gradient calculation?

Tracing One Backward Pass

Trace what happens after a network has produced a prediction that differs from the actual output.

Locate the final error: The backward pass begins only after the final output is available and its error has been determined.

Start at the later layer: The procedure begins propagating error information from the output side of the network.

Use local information: The computation uses the next layer's error quantity, the connecting weights, and the activation derivative.

Compute an edge gradient: The chain rule connects the error information and local derivative information so the gradient for a connection can be calculated.

Continue toward earlier layers: The same backward pattern continues from later layers toward earlier layers.

The backward pass turns the final error into gradients that describe how the model parameters contributed to that error.

Forward and Backward Compared

Forward passBackward pass
Computes weighted sums, activations, and the final output.Uses the chain rule to propagate error information and compute gradients.
Moves from the input through later layers.Moves from later layers toward earlier layers.
Must occur before the final error is available.Begins after the final output and its error are available.

Checking the Procedure

A correct trace preserves a specific order: forward computation, final error, backward computation, and edge-gradient calculation. To find a divergence, inspect the phase order rather than looking only at an individual calculation.

completesmakes possiblestartscomputesif skipped or reorderedif unavailableif direction changesForward computationFinal outputFinal errorBackward computationEdge-gradientcalculationSequence divergence
What should happen at each stage, and where can a procedure branch away from the expected sequence?
  • Starting the backward pass before the forward pass has produced the final output.

    The final error depends on having the network output available.

    Fix: Complete the forward computation first, then determine the final error.

  • Moving error information from earlier layers toward later layers during the backward pass.

    The expected backward direction is from later layers toward earlier layers.

    Fix: Begin with the final error and trace the error information backward through the layers.

  • Treating the chain rule as the forward computation.

    The forward pass computes the output; the chain rule is used during the backward pass to compute gradients.

    Fix: Separate prediction-making from gradient computation.

  • Skipping the edge-gradient calculation.

    Computing gradients for model parameters is the central task of backpropagation.

    Fix: Check that the backward trace includes the connecting weights, activation derivative, and edge-gradient calculation.

Practice the Trace

MEDIUM

A learner writes this sequence: final error, forward weighted sums, backward computation, activations, edge-gradient calculation. Identify the first point where the sequence diverges from the expected backpropagation order, and rewrite the sequence in the correct phase order.

Hints
  • Ask whether the final error can be determined before the final output exists.
  • Separate all forward-pass actions from all backward-pass actions.
  • Remember that edge-gradient calculation belongs after the backward computation begins.

Checking the Practice Sequence

Find the first divergence in the sequence: final error, forward weighted sums, backward computation, activations, edge-gradient calculation.

Check the first step: The sequence begins with the final error, but the forward computation has not yet produced the final output.

Identify the divergence: The first step is the divergence because the forward pass must come before the final error.

Restore the forward phase: The weighted sums and activations must be computed through the layers so the final output becomes available.

Begin the backward phase: After the final error is available, the backward computation can move from later layers toward earlier layers.

Finish with gradients: The procedure then calculates gradients for the connecting edges.

A correct high-level order is forward weighted sums and activations, final output, final error, backward computation, and edge-gradient calculation.

Summary

  1. The forward pass computes weighted sums, activations, and the final output.
  2. The final error must be available before the backward pass begins.
  3. The backward pass uses the chain rule to move error information from later layers toward earlier layers.
  4. Connecting weights and activation derivatives contribute to edge-gradient calculations.
  5. The expected order is forward computation, final error, backward computation, and edge-gradient calculation.

Key Takeaways

  • The forward pass transforms an input through weighted sums and activations until it produces a final output.
  • The backward pass begins with the error between the prediction and actual output.
  • The chain rule organizes how error information moves backward and how gradients are computed.
  • A reliable trace preserves the order from forward computation to final error, backward computation, and edge-gradient calculation.
  • To diagnose a divergence, check both the phase order and the direction of information flow.