Concepts / Differentiable Activation Functions

Differentiable Activation Functions

Backpropagation alternates forward and backward passes through an ANN with hidden layers.

  • Programming

Why One Pass Is Not Enough

Training an artificial neural network with hidden layers requires more than sending an input through the network once. The network first calculates the activations of its units. It then works backward to determine a partial derivative for each weight. Backpropagation organizes these two activities into alternating forward and backward passes.

The forward pass determines what the network produces for the current input and current network state. The backward pass determines how the weights are represented in the gradient estimate used during training.

What do you think happens?

After a forward pass has computed the network's unit activations, what does backpropagation do next?

  • It repeats the same forward computation only
  • It computes a partial derivative for each weight
  • It removes the hidden layers
  • It stops because the input has already passed through the network
Reveal answer

Answer: It computes a partial derivative for each weight.

The backward pass does not simply recompute the forward activations. Its task is to compute a partial derivative for each weight, and the collected derivatives form a gradient estimate.

Forward Computation Through the Network

A forward pass begins with the current activations of the input units. Computation then proceeds through the network toward later units. Each unit's activation is determined from the activations available before it. In this direction, information moves from the input side toward the later units, producing the activations associated with the current input and the current network state.

forward computationforward computationforward computationInput activationscurrent valuesHidden-unitactivationscomputed from earlieractivationsLater-unitactivationscomputed from availableactivationsOutput activationresult for the currentinput
What values are computed as input activations move through the network?

Tracing a Forward Pass

Describe what happens when an input is sent through an ANN with hidden layers during a forward pass.

Begin with the input side: The pass starts with the current activations of the input units.

Compute hidden-unit activations: The network determines each hidden unit's activation from the activations available before that unit.

Continue toward later units: The same forward direction continues through the network, from the input side toward later units.

Record the resulting activations: The completed pass tells the network what activations result from the current input and current network state.

A forward pass computes the network's activations in input-to-later-unit order. It provides the activations that the subsequent backward pass works with.

Backward Derivative Computation

After the forward pass, backpropagation performs a backward pass. The backward pass has a different task: it does not recompute the forward activations. Instead, it computes a partial derivative for each weight. The complete collection of these partial derivatives forms an estimate of the true gradient.

For an ANN with hidden layers, this derivative computation must account for weights throughout the network, not only for weights nearest the output. This is why the backward stage is essential to training: it produces derivative information for every weight, including weights associated with hidden-layer computation.

backward computationcontinues through the networkcollect derivativesOutput activationsafter the forward passLater-weightderivativespartial derivativesHidden-weightderivativespartial derivativesGradient estimatecollected derivatives
How does the backward computation move from the later units toward derivative information for weights throughout the network?
Forward passBackward pass
Computes unit activationsComputes a partial derivative for each weight
Starts with current input-unit activationsFollows the forward pass
Moves from the input side toward later unitsProduces derivative information for the weights
Describes the activations resulting from the current input and network stateCollects derivatives into a gradient estimate

The two passes have different jobs but work together during training.

The Training Cycle

Backpropagation is best understood as a cycle rather than as a single one-way calculation. First, a forward pass computes unit activations. After that pass, a backward pass computes a partial derivative for every weight. The collected derivatives form a gradient estimate that the training process can use. While the network is being trained, these two directions alternate.

startthencollect derivativesused during trainingnext cycleCurrent networkstatecurrent weights andactivationsForward passunit activationsBackward passweight derivativesGradient estimatecollected derivativesTraining processuses the estimate
What happens after a forward pass produces activations, and what leads into the next training cycle?

Following One Training Cycle

Trace the information produced in one alternating forward-and-backward cycle.

Forward stage: The network starts from the current input-unit activations and computes activations through the later units.

Backward stage: Once the forward activations have been computed, the network performs a backward pass rather than repeating the forward calculation.

Derivative stage: The backward pass computes a partial derivative for each weight.

Gradient stage: The collected partial derivatives form an estimate of the true gradient for use by the training process.

Repeat: Because training uses alternating passes, another cycle can begin from the network's current state.

The cycle connects activations to derivative information: forward computation produces the activations, and backward computation turns the network's weight-related information into a gradient estimate.

Common Misunderstandings

  • Treating backpropagation as only a forward calculation

    Backpropagation includes a backward pass that computes a partial derivative for each weight.

    Fix: Describe the process as a forward pass followed by a backward pass.

  • Saying that the backward pass recomputes the forward activations

    The source distinguishes the tasks: the forward pass computes activations, while the backward pass computes partial derivatives.

    Fix: Associate the backward pass with derivative computation rather than with recomputing forward activations.

  • Ignoring hidden-layer weights

    The backward pass computes a partial derivative for each weight, including weights associated with hidden-layer computation.

    Fix: Track derivative computation across the network rather than stopping at the output side.

  • Calling one forward-and-backward sequence a complete explanation of training

    Backpropagation is organized as an alternating cycle used while the network is being trained.

    Fix: Explain that forward and backward passes alternate during training.

When explaining a training step, name the information produced at each stage. Say activations for the forward pass, partial derivatives for the backward pass, and gradient estimate for the collected derivatives.

Practice Check

MEDIUM

An ANN with hidden layers has completed a forward pass. Explain, in order, what the backward pass computes and how its results support the next stage of training.

Hints
  • Start by naming the unit-level information computed during the forward pass.
  • State what is computed for each weight during the backward pass.
  • Explain what the collection of those derivatives represents.
EASY

Classify each statement as describing the forward pass or the backward pass: computing unit activations; starting from input-unit activations; computing a partial derivative for each weight; collecting derivatives into a gradient estimate.

Hints
  • The forward pass moves from the input side toward later units.
  • The backward pass follows the forward pass and focuses on derivatives.

Key Takeaways

  1. An ANN with hidden layers needs more than a single forward calculation during training.
  2. The forward pass starts from current input-unit activations and computes activations through later units.
  3. The backward pass computes a partial derivative for each weight rather than recomputing the forward activations.
  4. The collected partial derivatives form an estimate of the true gradient.
  5. Backpropagation supports training by alternating forward and backward passes.

Key Takeaways

  • Forward passes compute unit activations from the current input-side activations.
  • Backward passes compute partial derivatives for the network's weights.
  • The collected derivatives form a gradient estimate used by the training process.
  • Backpropagation is a repeating cycle of alternating forward and backward passes.