Differentiable Activation Functions
Backpropagation alternates forward and backward passes through an ANN with hidden layers.
Why One Pass Is Not Enough
Training an artificial neural network with hidden layers requires more than sending an input through the network once. The network first calculates the activations of its units. It then works backward to determine a partial derivative for each weight. Backpropagation organizes these two activities into alternating forward and backward passes.
The forward pass determines what the network produces for the current input and current network state. The backward pass determines how the weights are represented in the gradient estimate used during training.
What do you think happens?
After a forward pass has computed the network's unit activations, what does backpropagation do next?
Reveal answer
Answer: It computes a partial derivative for each weight.
The backward pass does not simply recompute the forward activations. Its task is to compute a partial derivative for each weight, and the collected derivatives form a gradient estimate.
Forward Computation Through the Network
A forward pass begins with the current activations of the input units. Computation then proceeds through the network toward later units. Each unit's activation is determined from the activations available before it. In this direction, information moves from the input side toward the later units, producing the activations associated with the current input and the current network state.
Tracing a Forward Pass
Describe what happens when an input is sent through an ANN with hidden layers during a forward pass.
Begin with the input side: The pass starts with the current activations of the input units.
Compute hidden-unit activations: The network determines each hidden unit's activation from the activations available before that unit.
Continue toward later units: The same forward direction continues through the network, from the input side toward later units.
Record the resulting activations: The completed pass tells the network what activations result from the current input and current network state.
A forward pass computes the network's activations in input-to-later-unit order. It provides the activations that the subsequent backward pass works with.
Backward Derivative Computation
After the forward pass, backpropagation performs a backward pass. The backward pass has a different task: it does not recompute the forward activations. Instead, it computes a partial derivative for each weight. The complete collection of these partial derivatives forms an estimate of the true gradient.
For an ANN with hidden layers, this derivative computation must account for weights throughout the network, not only for weights nearest the output. This is why the backward stage is essential to training: it produces derivative information for every weight, including weights associated with hidden-layer computation.
| Forward pass | Backward pass |
|---|---|
| Computes unit activations | Computes a partial derivative for each weight |
| Starts with current input-unit activations | Follows the forward pass |
| Moves from the input side toward later units | Produces derivative information for the weights |
| Describes the activations resulting from the current input and network state | Collects derivatives into a gradient estimate |
The two passes have different jobs but work together during training.
The Training Cycle
Backpropagation is best understood as a cycle rather than as a single one-way calculation. First, a forward pass computes unit activations. After that pass, a backward pass computes a partial derivative for every weight. The collected derivatives form a gradient estimate that the training process can use. While the network is being trained, these two directions alternate.
Following One Training Cycle
Trace the information produced in one alternating forward-and-backward cycle.
Forward stage: The network starts from the current input-unit activations and computes activations through the later units.
Backward stage: Once the forward activations have been computed, the network performs a backward pass rather than repeating the forward calculation.
Derivative stage: The backward pass computes a partial derivative for each weight.
Gradient stage: The collected partial derivatives form an estimate of the true gradient for use by the training process.
Repeat: Because training uses alternating passes, another cycle can begin from the network's current state.
The cycle connects activations to derivative information: forward computation produces the activations, and backward computation turns the network's weight-related information into a gradient estimate.
Common Misunderstandings
Treating backpropagation as only a forward calculation
Backpropagation includes a backward pass that computes a partial derivative for each weight.
Fix:
Describe the process as a forward pass followed by a backward pass.Saying that the backward pass recomputes the forward activations
The source distinguishes the tasks: the forward pass computes activations, while the backward pass computes partial derivatives.
Fix:
Associate the backward pass with derivative computation rather than with recomputing forward activations.Ignoring hidden-layer weights
The backward pass computes a partial derivative for each weight, including weights associated with hidden-layer computation.
Fix:
Track derivative computation across the network rather than stopping at the output side.Calling one forward-and-backward sequence a complete explanation of training
Backpropagation is organized as an alternating cycle used while the network is being trained.
Fix:
Explain that forward and backward passes alternate during training.
When explaining a training step, name the information produced at each stage. Say activations for the forward pass, partial derivatives for the backward pass, and gradient estimate for the collected derivatives.
Practice Check
An ANN with hidden layers has completed a forward pass. Explain, in order, what the backward pass computes and how its results support the next stage of training.
Hints
- Start by naming the unit-level information computed during the forward pass.
- State what is computed for each weight during the backward pass.
- Explain what the collection of those derivatives represents.
Classify each statement as describing the forward pass or the backward pass: computing unit activations; starting from input-unit activations; computing a partial derivative for each weight; collecting derivatives into a gradient estimate.
Hints
- The forward pass moves from the input side toward later units.
- The backward pass follows the forward pass and focuses on derivatives.
Key Takeaways
- An ANN with hidden layers needs more than a single forward calculation during training.
- The forward pass starts from current input-unit activations and computes activations through later units.
- The backward pass computes a partial derivative for each weight rather than recomputing the forward activations.
- The collected partial derivatives form an estimate of the true gradient.
- Backpropagation supports training by alternating forward and backward passes.
Key Takeaways
- Forward passes compute unit activations from the current input-side activations.
- Backward passes compute partial derivatives for the network's weights.
- The collected derivatives form a gradient estimate used by the training process.
- Backpropagation is a repeating cycle of alternating forward and backward passes.