Concepts / Backpropagation

Backpropagation

Tensors provide the multi-dimensional representations used by neural networks.

  • Programming

From Data to Learning

A neural network performs computation by passing data through connected operations. The data is represented as tensors, and the final computation produces a loss value. Backpropagation begins with that final loss and moves backward through the connected operations to calculate gradient values. Those gradients provide information for optimizing the network's parameters.

Backpropagation is an algorithm for computing the gradient values of a neural network.

Tensor Representations

Tensors are the multi-dimensional arrays used to represent neural-network data. The word tensor describes the dimensional organization of the data: a scalar has no dimensions, a vector has one dimension, a matrix has two dimensions, and a higher-dimensional tensor has more than two dimensions. Tensor operations then manipulate the values or the organization of these representations.

RepresentationDimensionsRole in the neural-network data model
ScalarNo dimensionsA single value
VectorOne dimensionA one-dimensional collection of values
MatrixTwo dimensionsA two-dimensional organization of values
Higher-dimensional tensorMore than two dimensionsA multi-dimensional organization of values

Tensor representations differ by the number of dimensions used to organize their values.

Scalarone valueVectorone dimensionMatrixtwo dimensionsHigher-dimensionaltensormore than two dimensions
What do scalars, vectors, matrices, and higher-dimensional tensors contain, and how do their dimensions organize neural-network data?

Tensor operations are the computational machinery of a neural network. Element-wise operations work across corresponding values. Broadcasting supports computation across dimensions when an operation uses differently organized tensors. Tensor dot operations combine tensor values through a dot operation. Reshaping changes the organization of values. These roles should be kept distinct: an operation may change values, organization, or both.

valuesdimensionsvaluesorganizationtransformed valuesdimension-aware computationcombined valuesnew organizationInput tensorvalues and organizationElement-wiseoperationcorresponding valuesOutput tensortransformed representationBroadcastingacross dimensionsTensor dotcombine tensor valuesReshapingchange organization
How can tensor operations change values, shape, or arrangement?

Forward Computation

A neural network can be viewed as a sequence of operations. The output of one operation becomes the input to another operation. As this forward computation proceeds, tensor operations manipulate the network's representations and produce a final result. The final result produces a loss value, which provides the starting point for the backward computation.

forward dataforward resultproducesbackward gradient informationmoves toward earlier operationsInput tensorConnected operationsFinal resultLossGradient values
What data moves forward to produce the prediction, and what gradient information moves backward from the loss?

Tracing a Connected Computation

Trace the direction of information through a neural network viewed as connected operations.

Represent: The network represents its data as tensors, which may be scalars, vectors, matrices, or higher-dimensional tensors.

Compute: Each operation receives input from an earlier operation and produces output for a later operation.

Measure: The final result produces a loss value.

Reverse: Backpropagation starts with that final loss and follows the connected operations backward.

The forward route produces the loss; the reverse route computes gradient values that describe parameter contributions to that loss.

Reverse Gradient Trace

Backpropagation starts with the final loss and moves backward from the top layers toward the bottom layers. At each connected operation, it uses derivative information from the later part of the network together with the local derivative of the current operation. This backward traversal follows the reverse of the forward nesting order.

backward derivative informationcombined derivative informationcombined derivative informationparameter gradientsLossstarting pointTop operationlocal derivativeEarlier operationlocal derivativeBottom operationparameter contributionGradient valuescomputed backward
How does backpropagation move from the final loss backward through each connected operation to compute gradients?

The Chain Rule

A network is built from connected operations, so a parameter can affect the final loss through more than one operation. The chain rule combines the local derivatives of those connected operations. Backpropagation uses this chained derivative information to determine how each parameter contributed to the final loss.

influencesconnected computationcontributes toreverse chainlocal derivativelocal derivativeParameterOperation 1local derivativeOperation 2local derivativeLossParametercontributioncombined derivativeinformation
How are local derivatives combined along a sequence of operations to determine how one parameter affects the final loss?

Following One Parameter's Influence

Explain how the contribution of one parameter is determined when it affects a loss through connected operations.

Locate the parameter: Start with the parameter inside the forward sequence of tensor operations.

Reach the loss: Follow the forward connections from that operation to the final loss.

Reverse the route: Start at the final loss and move backward through the same connected operations in reverse order.

Combine local information: At every operation, combine the derivative information arriving from later operations with the local derivative of the current operation.

The resulting gradient value represents how that parameter contributed to the final loss.

Gradient-Based Updates

Computation alone does not explain learning. Learning uses gradient-based optimization to optimize the network's parameters and minimize the loss function. The gradient is described as the derivative of a tensor operation, so gradient values provide derivative information about the computation. Stochastic gradient descent is one method used in this optimization process.

used byproducesstarting pointcomputesderivative informationoptimizesParametersTensor computationLossBackpropagationGradientsStochastic gradientdescentOptimized parameters
How does the loss produce gradients, and how do those gradients change network parameters during one training step?

The relationship between backpropagation and stochastic gradient descent is important: backpropagation chains derivatives so that the effects of tensor operations can contribute to parameter updates. Backpropagation computes the gradient information; gradient-based optimization uses that information to optimize parameters and minimize loss.

Common Tracing Mistakes

  • Treating backpropagation as another forward computation.

    Backpropagation begins with the final loss and moves backward through the connected operations.

    Fix: Start at the loss, inspect the top operation, and continue toward earlier operations.

  • Ignoring the connected structure of the network.

    A parameter's effect can pass through later connected operations before reaching the loss.

    Fix: Use the chain rule to combine local derivatives along the connected route.

  • Confusing tensor computation with learning.

    Tensor operations provide the computation, but learning uses gradient-based optimization to optimize parameters and minimize loss.

    Fix: Separate the forward computation, gradient calculation, and optimization step conceptually.

  • Treating all tensor operations as having the same role.

    Element-wise operations, broadcasting, tensor dot operations, and reshaping manipulate values and organization in different ways.

    Fix: Identify whether the operation works on corresponding values, supports computation across dimensions, combines tensor values, or changes organization.

  • Checking a backward calculation from the wrong end.

    The backward route is defined by starting at the final loss and moving toward earlier operations.

    Fix: Find the first operation in the backward route where later derivative information was not combined with the current local derivative according to the chain-rule structure.

Practice the Backward Route

MEDIUM

A network is described as three connected operations followed by a loss. Explain, in order, how you would trace backpropagation to determine the contribution of a parameter located in the first operation.

Hints
  • Begin with the final loss rather than the parameter.
  • Move through the operations in the reverse of their forward nesting order.
  • At each operation, identify the later derivative information and the current local derivative.
  • End by describing the resulting gradient as information about the parameter's contribution to the loss.

What do you think happens?

A parameter appears in an early operation, while the loss is produced by later operations. Which direction should you follow to determine the parameter's gradient contribution?

  • From the parameter forward to the loss only
  • From the loss backward through the connected operations
  • From the tensor representation directly to stochastic gradient descent
  • From the loss to the input without inspecting intermediate operations
Reveal answer

Answer: From the loss backward through the connected operations

Backpropagation starts with the final loss, follows the reverse of the forward nesting order, and combines later derivative information with each operation's local derivative.

Key Takeaways

  1. Tensors are multi-dimensional arrays used to represent neural-network data; they may be scalars, vectors, matrices, or higher-dimensional tensors.
  2. Tensor operations manipulate values, combine values, support computation across dimensions, or change organization.
  3. A neural network's connected operations produce a final loss during the forward computation.
  4. Backpropagation starts with the final loss and moves backward through the connected operations.
  5. The chain rule combines local derivatives so gradient values can describe each parameter's contribution to the loss.
  6. Stochastic gradient descent is a gradient-based optimization method that uses derivative information to optimize parameters and minimize loss.

Key Takeaways

  • Neural networks represent data and computations with tensors.
  • Tensor operations can manipulate values, combine values, work across dimensions, or change organization.
  • Backpropagation computes gradient values by starting at the final loss and moving backward through connected operations.
  • The chain rule combines local derivatives to determine how parameters contributed to the loss.
  • Gradient-based optimization, including stochastic gradient descent, uses this derivative information to optimize parameters and minimize loss.