Concepts / SGD and Backpropagation for Neural Networks

SGD and Backpropagation for Neural Networks

Backpropagation calculates the gradient of a loss with respect to network weights for an example.

  • Programming

The Two-Direction Calculation

Backpropagation calculates the gradient of a loss with respect to a neural network's weights for one example (x, y). It does this by evaluating the network in two directions. The forward pass moves from the input toward the output and records the quantities produced by each layer. The backward pass starts at the output and moves toward earlier layers, combining the information needed to compute derivatives.

forward dataforward datapredictionbackward errorbackward erroredge derivativesInput xo_0 = xWeight gradientsone gradient per relevantedgeLayer 1a_1, o_1Error at Layer 1backward calculationLayer 2a_2, o_2Error at Layer 2backward calculationLossprediction compared with y
How do data and gradient information move through a neural network in opposite directions?

The forward pass is necessary because the backward calculation uses forward quantities, including each layer's pre-activation and output. The backward pass is necessary because it carries loss sensitivity from the final layer toward earlier layers.

Forward Pass State

Think of the network as a sequence of composed functions. The input is the initial output, written as o_0 = x. At layer t, the network first forms a pre-activation vector by multiplying the layer's weight matrix by the previous output: a_t = W_(t-1)o_(t-1). It then applies the activation function to produce o_t = σ(a_t). The forward pass repeats this process from the input toward the final output.

Tracing One Layer's Forward Quantities

Identify the quantities that must be available when the layer connecting o_(t-1) to o_t is evaluated.

Receive the previous output: The layer receives o_(t-1), which was produced by the preceding layer.

Form the pre-activation: The layer computes a_t = W_(t-1)o_(t-1).

Apply the activation: The layer computes o_t = σ(a_t).

Retain the forward state: The pre-activation a_t and output o_(t-1) remain relevant to the later backward calculation.

The forward pass produces the intermediate values that the backward pass needs, not only the final prediction.

Backward Error Signals

The backward pass begins at the final layer. For the squared loss used in the source material, the final error vector is δ_T = o_T - y. Earlier error vectors are then computed by combining the next layer's weights, the next layer's error vector, and the derivative of the next layer's activation. This recursion moves the loss information toward the first layer.

local derivativenext-layer derivativelater effectEarly weighteffect begins locallyEarlier outputo_tLater outputsuccessive transformationLossfinal effect
How are derivatives from successive layers connected when determining how a loss changes with respect to an earlier layer's weights?

The chain rule explains why an early weight cannot be differentiated by looking only at its immediate layer. A change in that weight changes the next layer's output, which changes later outputs, which changes the loss. The derivative of the full network is therefore assembled from the derivatives of these successive transformations, in the appropriate order.

δ_T = o_T - y

Jacobian View

A Jacobian organizes partial derivatives by output component and input variable. Its shape records how many outputs depend on how many inputs. This makes it useful for describing the derivative of a layer whose input and output are vectors rather than single numbers.

input columninput columnoutput rowoutput rowInput variable 1Output component 1partial derivatives acrossinputsJacobianrows: outputs, columns:inputsInput variable 2Output component 2partial derivatives acrossinputs
How does a Jacobian organize the effects of multiple input variables on multiple output components?

Reading Two Jacobian Cases

Compare the Jacobian description for a linear function with the Jacobian description for an element-wise sigmoid transformation.

Linear function: For f(w) = Aw, the Jacobian is A itself.

Element-wise sigmoid: For an element-wise sigmoid transformation, the Jacobian is diagonal.

Interpret the diagonal: The diagonal entries contain the individual sigmoid derivatives.

Interpret the off-diagonal entries: The off-diagonal entries are zero for this element-wise transformation.

The Jacobian is not merely a generic matrix label: its shape and entries describe which output components respond to which input variables.

An Edge Gradient

Consider the weights connecting layer t - 1 to layer t. The layer computes a_t = W_(t-1)o_(t-1), followed by o_t = σ(a_t). When differentiating with respect to the weights in this particular matrix, the earlier output o_(t-1) is treated as fixed. The derivative for an edge is then assembled from three local pieces: the error at the destination, the activation derivative at the destination, and the output at the source.

source factorerror factorlocal derivative factorSource outputo_(t-1)Edge gradientderivative for one weightDestination errorearlier error vectorActivationderivativeat destinationpre-activation
How do the upstream error, destination activation derivative, and source activation combine for one network edge?

This factorization is a debugging aid. If the source output, destination activation derivative, and relevant error vector are correct, inspect the final assembly of the edge derivative rather than immediately blaming the forward pass or the recursive backward calculation.

Gradient Discrepancy Diagnosis

A gradient mismatch should be investigated as a sequence of checkpoints. First inspect the forward pre-activations a_t and outputs o_t. Then inspect the final error δ_T and the recursively computed earlier error vectors δ_t. Finally inspect the derivative assembled for each edge. This order follows the dependency structure of backpropagation and helps distinguish a forward error from a backward error or an edge-calculation error.

check firstincorrectcorrectincorrectcorrectincorrectGradient mismatchForward quantitiesa_t and o_tForward-pass issueFinal errorδ_TBackward-pass issueerror recursionEdge derivativefinal assemblyEdge-gradient issue
How can a gradient discrepancy be localized to the forward pass, backward pass, or final edge-gradient calculation?
  • Checking only the final gradient value

    The gradient depends on forward pre-activations, forward outputs, error vectors, and the final edge assembly.

    Fix: Inspect the checkpoints in dependency order: a_t, o_t, δ_T, earlier δ_t values, and then the edge derivative.

  • Ignoring stored forward quantities

    The edge derivative needs the source output and the destination pre-activation.

    Fix: Retain the forward values needed by the later backward calculation.

  • Treating an early weight as affecting only its immediate layer

    A change in an early weight propagates through successive layer transformations before affecting the loss.

    Fix: Use the chain rule to combine the local derivatives across the composed functions.

SGD Connection

Backpropagation supplies the gradient of the loss with respect to the network's weights for one example. That gradient is the information an SGD procedure uses when working with the weights. The key boundary to remember is that backpropagation is the gradient-calculation procedure: it traces how the loss depends on the weights and produces the corresponding derivatives.

backpropagation evaluates dependenceSGD uses gradient informationWeightcurrent network parameterWeight gradientfrom one exampleWeightparameter after update step
How does the gradient calculated for one example connect to the weight state used by an SGD update?

Practice Check

MEDIUM

A gradient for an early-layer edge is incorrect. The forward pre-activations and outputs are correct, and the final error vector is correct. The earlier error vectors are also correct. Which part should you inspect next, and which three quantities should you verify before inspecting the final assembly?

Hints
  • Use the dependency order of the backpropagation checkpoints.
  • Focus on the local calculation for the edge's destination and source.

What do you think happens?

Which checkpoint is the most likely next place to inspect?

  • The input example only
  • The source output, destination activation derivative, and relevant error vector
  • The network architecture name
  • The final prediction only
Reveal answer

Answer: The source output, destination activation derivative, and relevant error vector

Once the forward quantities and recursively computed error vectors are correct, the remaining issue may be in the local edge-gradient inputs or in how the edge derivative was assembled.

Key Takeaways

  1. The forward pass computes each layer's pre-activation and output, while the backward pass computes error vectors and partial derivatives.
  2. The chain rule connects local derivatives across the successive functions that compose a neural network.
  3. A Jacobian organizes partial derivatives by output component and input variable; its shape records the relationship between the numbers of outputs and inputs.
  4. An edge gradient depends on the destination error, the destination activation derivative, and the source output.
  5. When a gradient is unexpected, inspect forward quantities, error vectors, and final edge-gradient assembly as separate checkpoints.

Key Takeaways

  • Backpropagation calculates the gradient of a loss with respect to network weights for one example.
  • The forward pass creates the intermediate values that the backward pass needs.
  • The backward pass applies the chain rule from the final layer toward earlier layers.
  • Jacobians organize the partial derivatives of vector-valued layer transformations.
  • Gradient debugging should separate forward-state errors, backward-error errors, and final edge-gradient assembly errors.