Neural Network Loss Functions
Backpropagation calculates, rather than directly applies, the gradient of the loss with respect to network weights.
From Prediction to Weight Updates
A neural network has a loss that depends on its weights. The network first transforms an input into a prediction, and the loss measures the result. To change the weights in a useful direction, we need to know how the loss changes when each individual weight changes. Backpropagation calculates this information: it calculates the gradient of the loss with respect to the weights, but it does not itself apply the weight update.
Forward Values Through Layers
Consider a network divided into layers V0 through VT. The weights connecting layer Vt-1 to layer Vt are represented by a matrix Wt-1. The incoming values at a layer are collected in at, and the neuron outputs are collected in ot. The layer first computes at = Wt-1 ot-1 and then applies the activation function element by element to obtain ot = σ(at). These forward quantities are calculated from the input side toward the output side.
Tracking One Layer's Forward Quantities
Describe the forward computation at layer Vt using the notation from the network model.
Collect incoming values: The outputs from the preceding layer are represented by ot-1.
Apply the weight matrix: The layer computes the pre-activation values at = Wt-1 ot-1.
Apply the activation: The activation function is applied element by element, producing ot = σ(at).
Store the forward quantities: The values at and ot are retained because the backward calculation uses the forward quantities while moving in the opposite direction.
The forward pass records the values computed from the input side toward the output side, preparing information for backpropagation.
Partial Derivatives and Weight Sensitivity
A partial derivative describes how a function changes when one variable changes while the other variables are held constant. In a neural network, the loss has one partial derivative for each weight associated with an edge in the network. Each derivative therefore answers a specific question: how does the loss change if this particular weight changes while the other weights remain fixed?
The gradient collects all of these partial derivatives. If the network has many weights, the gradient has one corresponding component for each weight. The component associated with a particular weight reports the loss sensitivity for that weight; it is not a general score detached from the network's individual connections.
Jacobians for Vector Outputs
A function can produce several outputs at once. When a function maps Rⁿ to Rᵐ, its derivatives are organized in a Jacobian matrix with m rows and n columns. The entry in row i and column j is the partial derivative of output i with respect to input variable j. Rows therefore correspond to output components, while columns correspond to input components.
When there is only one output, the Jacobian is the gradient represented as a row vector. In a multilayer network, Jacobians describe how one layer's outputs respond to another layer's inputs, making them the local derivative objects used by the chain rule.
The Chain Rule Across Layers
A neural network composes many functions. A weight matrix produces pre-activation values, an activation function produces neuron outputs, later layers transform those outputs, and the final output determines the loss. The chain rule connects the derivative of the final loss to the derivatives of each intermediate transformation.
The order of the derivative operations matters because each derivative maps changes from one representation to the next. Applying the chain rule means combining local derivative information in the correct order. Backpropagation organizes this repeated multiplication recursively so that sensitivity arriving at one layer can be reused when calculating the sensitivity for the preceding layer.
Following One Weight's Influence
Trace conceptually how a change in a weight near the input side can influence the final loss.
Weight to pre-activation: The weight contributes to a layer's pre-activation values through the weight-matrix transformation.
Pre-activation to output: The activation function transforms the pre-activation values into the layer outputs.
Output through later layers: Later layers transform those outputs and pass the effect toward the final output.
Final output to loss: The final output determines the loss, so the original weight can affect the loss through every intermediate transformation.
Combine local derivatives: The chain rule combines the local derivative at each link into the derivative of the complete network with respect to that weight.
The derivative for the weight is obtained by composing the local sensitivities along the entire path from that weight to the loss.
Forward Pass and Backward Pass
Backpropagation has two directions of computation. During the forward pass, the network calculates the values at and ot from the input side toward the output side. During the backward pass, it uses those stored forward quantities while moving from the output side toward the input side. The forward values describe what the network computed; the backward sensitivities describe how strongly the later loss depends on those computed outputs.
Let ℓt be the loss of the subnetwork beginning at layer Vt, viewed as a function of the outputs ot. The backward sensitivity is defined as δt = Jot(ℓt). This vector records how the subnetwork loss changes with respect to the outputs at that layer.
The backward quantities are not another set of forward activations. They are sensitivities. A sensitivity tells us how strongly a later loss depends on an output currently being considered.
Backward Recursion and Weight Gradients
At the final layer, the subnetwork loss is the loss function itself. For the squared loss Δ(u, y) = 1/2 ||u − y||², the derivative with respect to u is u − y. Therefore, at the final layer, δT = oT − y. This supplies the starting backward signal.
For an earlier layer, the chain rule expresses δt in terms of the sensitivity at the next layer, the derivative of the activation, and the transformation connecting the layers. The exact matrix arrangement follows from the Jacobians of those transformations. Repeating this step moves the sensitivity from the top layer toward the bottom layer.
Tracing the Backward Signal
Explain how the backward recursion moves from the final layer to an earlier layer and contributes to the weight gradient.
Start at the final layer: For the squared loss, the final sensitivity is δT = oT − y.
Move to the preceding layer: Use the next layer's sensitivity together with the activation derivative and the connecting transformation.
Repeat toward the input: Apply the same chain-rule reasoning repeatedly so each earlier layer receives a sensitivity.
Combine with forward information: The stored forward quantities and the backward sensitivities are combined to obtain derivatives for the weights.
The backward recursion produces the gradient of the loss with respect to the network weights by repeatedly applying local derivative relationships from the output side toward the input side.
Common Reasoning Mistakes
Treating the gradient as one undifferentiated number.
The gradient collects one partial derivative for each network weight.
Fix:
Track every gradient component as the derivative of the loss with respect to its corresponding weight.Confusing forward values with backward sensitivities.
The forward quantities at and ot describe computed values, while δt describes how the loss depends on those outputs.
Fix:
Label forward quantities as values and δ quantities as sensitivities.Ignoring the order of derivative composition.
Each derivative maps changes from one representation to the next, so the order matters.
Fix:
Follow the path from the weight through pre-activations, activations, later layers, and finally the loss.Assuming backpropagation directly changes the weights.
Backpropagation calculates the gradient; it does not directly apply it.
Fix:
Describe gradient calculation and weight updating as separate operations.Treating a Jacobian as a scalar derivative.
For a function from Rⁿ to Rᵐ, the Jacobian contains the partial derivative for every output-input pair.
Fix:
Use rows for output components and columns for input components.
Practice Check
A network contains several layers. Explain, in order, what information is computed during the forward pass, where the initial backward sensitivity comes from for the squared loss, and how the chain rule carries that sensitivity toward earlier layers. End by explaining why the resulting gradient has one component for each network weight.
Hints
- Begin with the quantities at and ot.
- Use the final-layer relation δT = oT − y for the squared loss.
- Distinguish the local derivatives at each layer from the complete derivative of the loss with respect to a weight.
- Connect each gradient component to one particular weight.
Key Takeaways
- A partial derivative measures how the loss changes when one weight varies while the other weights are held constant.
- The gradient collects one such partial derivative for every network weight.
- A Jacobian organizes derivatives of every output component with respect to every input component.
- The chain rule combines local derivatives across the transformations in a multilayer network.
- Backpropagation moves sensitivities backward, combines them with stored forward quantities, and calculates the gradient; a separate process can then use that gradient to change the weights.
Key Takeaways
- A partial derivative links one network weight to the way the loss changes.
- A gradient is a collection of these weight-specific partial derivatives.
- Jacobians arrange derivatives between vector-valued inputs and outputs.
- Backpropagation applies the chain rule recursively from the loss toward the input.
- Forward values and backward sensitivities work together to calculate the gradient of the loss with respect to network weights.