Backpropagation
Tensors provide the multi-dimensional representations used by neural networks.
From Data to Learning
A neural network performs computation by passing data through connected operations. The data is represented as tensors, and the final computation produces a loss value. Backpropagation begins with that final loss and moves backward through the connected operations to calculate gradient values. Those gradients provide information for optimizing the network's parameters.
Backpropagation is an algorithm for computing the gradient values of a neural network.
Tensor Representations
Tensors are the multi-dimensional arrays used to represent neural-network data. The word tensor describes the dimensional organization of the data: a scalar has no dimensions, a vector has one dimension, a matrix has two dimensions, and a higher-dimensional tensor has more than two dimensions. Tensor operations then manipulate the values or the organization of these representations.
| Representation | Dimensions | Role in the neural-network data model |
|---|---|---|
| Scalar | No dimensions | A single value |
| Vector | One dimension | A one-dimensional collection of values |
| Matrix | Two dimensions | A two-dimensional organization of values |
| Higher-dimensional tensor | More than two dimensions | A multi-dimensional organization of values |
Tensor representations differ by the number of dimensions used to organize their values.
Tensor operations are the computational machinery of a neural network. Element-wise operations work across corresponding values. Broadcasting supports computation across dimensions when an operation uses differently organized tensors. Tensor dot operations combine tensor values through a dot operation. Reshaping changes the organization of values. These roles should be kept distinct: an operation may change values, organization, or both.
Forward Computation
A neural network can be viewed as a sequence of operations. The output of one operation becomes the input to another operation. As this forward computation proceeds, tensor operations manipulate the network's representations and produce a final result. The final result produces a loss value, which provides the starting point for the backward computation.
Tracing a Connected Computation
Trace the direction of information through a neural network viewed as connected operations.
Represent: The network represents its data as tensors, which may be scalars, vectors, matrices, or higher-dimensional tensors.
Compute: Each operation receives input from an earlier operation and produces output for a later operation.
Measure: The final result produces a loss value.
Reverse: Backpropagation starts with that final loss and follows the connected operations backward.
The forward route produces the loss; the reverse route computes gradient values that describe parameter contributions to that loss.
Reverse Gradient Trace
Backpropagation starts with the final loss and moves backward from the top layers toward the bottom layers. At each connected operation, it uses derivative information from the later part of the network together with the local derivative of the current operation. This backward traversal follows the reverse of the forward nesting order.
The Chain Rule
A network is built from connected operations, so a parameter can affect the final loss through more than one operation. The chain rule combines the local derivatives of those connected operations. Backpropagation uses this chained derivative information to determine how each parameter contributed to the final loss.
Following One Parameter's Influence
Explain how the contribution of one parameter is determined when it affects a loss through connected operations.
Locate the parameter: Start with the parameter inside the forward sequence of tensor operations.
Reach the loss: Follow the forward connections from that operation to the final loss.
Reverse the route: Start at the final loss and move backward through the same connected operations in reverse order.
Combine local information: At every operation, combine the derivative information arriving from later operations with the local derivative of the current operation.
The resulting gradient value represents how that parameter contributed to the final loss.
Gradient-Based Updates
Computation alone does not explain learning. Learning uses gradient-based optimization to optimize the network's parameters and minimize the loss function. The gradient is described as the derivative of a tensor operation, so gradient values provide derivative information about the computation. Stochastic gradient descent is one method used in this optimization process.
The relationship between backpropagation and stochastic gradient descent is important: backpropagation chains derivatives so that the effects of tensor operations can contribute to parameter updates. Backpropagation computes the gradient information; gradient-based optimization uses that information to optimize parameters and minimize loss.
Common Tracing Mistakes
Treating backpropagation as another forward computation.
Backpropagation begins with the final loss and moves backward through the connected operations.
Fix:
Start at the loss, inspect the top operation, and continue toward earlier operations.Ignoring the connected structure of the network.
A parameter's effect can pass through later connected operations before reaching the loss.
Fix:
Use the chain rule to combine local derivatives along the connected route.Confusing tensor computation with learning.
Tensor operations provide the computation, but learning uses gradient-based optimization to optimize parameters and minimize loss.
Fix:
Separate the forward computation, gradient calculation, and optimization step conceptually.Treating all tensor operations as having the same role.
Element-wise operations, broadcasting, tensor dot operations, and reshaping manipulate values and organization in different ways.
Fix:
Identify whether the operation works on corresponding values, supports computation across dimensions, combines tensor values, or changes organization.Checking a backward calculation from the wrong end.
The backward route is defined by starting at the final loss and moving toward earlier operations.
Fix:
Find the first operation in the backward route where later derivative information was not combined with the current local derivative according to the chain-rule structure.
Practice the Backward Route
A network is described as three connected operations followed by a loss. Explain, in order, how you would trace backpropagation to determine the contribution of a parameter located in the first operation.
Hints
- Begin with the final loss rather than the parameter.
- Move through the operations in the reverse of their forward nesting order.
- At each operation, identify the later derivative information and the current local derivative.
- End by describing the resulting gradient as information about the parameter's contribution to the loss.
What do you think happens?
A parameter appears in an early operation, while the loss is produced by later operations. Which direction should you follow to determine the parameter's gradient contribution?
Reveal answer
Answer: From the loss backward through the connected operations
Backpropagation starts with the final loss, follows the reverse of the forward nesting order, and combines later derivative information with each operation's local derivative.
Key Takeaways
- Tensors are multi-dimensional arrays used to represent neural-network data; they may be scalars, vectors, matrices, or higher-dimensional tensors.
- Tensor operations manipulate values, combine values, support computation across dimensions, or change organization.
- A neural network's connected operations produce a final loss during the forward computation.
- Backpropagation starts with the final loss and moves backward through the connected operations.
- The chain rule combines local derivatives so gradient values can describe each parameter's contribution to the loss.
- Stochastic gradient descent is a gradient-based optimization method that uses derivative information to optimize parameters and minimize loss.
Key Takeaways
- Neural networks represent data and computations with tensors.
- Tensor operations can manipulate values, combine values, work across dimensions, or change organization.
- Backpropagation computes gradient values by starting at the final loss and moving backward through connected operations.
- The chain rule combines local derivatives to determine how parameters contributed to the loss.
- Gradient-based optimization, including stochastic gradient descent, uses this derivative information to optimize parameters and minimize loss.