Concepts / Gradient Descent

Gradient Descent

Early neural networks were explored decades before modern deep learning, but their development was constrained by training difficulty.

  • Programming
Interactive lab

Try it: Gradient Descent

How gradient descent minimises a loss by repeatedly stepping against the gradient, and how the learning rate controls whether it converges, oscillates or diverges.

How it works

  1. The loss is f(x) = (x − 3)² + 1, lowest at x = 3.
  2. At the current x, compute the gradient f′(x) = 2(x − 3).
  3. Update x ← x − learning rate × gradient.
  4. Repeat. Small learning rates creep in, good ones converge quickly, rates above 1 overshoot further each step and diverge.

Default run (40 steps): Start at x = -3. Loss f(x) = (x − 3)² + 1 = 37. Learning rate 0.1. … Update 39: gradient f′ = 2(x − 3) = -0.0025; step = −0.1 × -0.0025 = 0.0002; new x = 2.999; loss 1 → 1. x is within 0.001 of the minimum at x = 3 — converged.

Simplified: One parameter and a perfectly smooth loss. Real models optimise millions of parameters on noisy losses.

Educational simulation

Loading the simulation…

Why Training Changed the Story

Neural networks were investigated in simple forms as early as the 1950s, decades before modern deep learning. The central idea existed, but the approach faced a practical obstacle: researchers did not yet have an efficient way to train large neural networks. In the mid-1980s, multiple people independently rediscovered Backpropagation. Combined with gradient-descent optimization, it provided a way to train chains of parametric operations. This changed the practical importance of neural networks because the challenge was no longer only imagining a network; it was finding useful values for its weights.

training obstacleresearch attention shifts1950searly neural-networkinvestigationmid-1980sBackpropagationrediscovered1990skernel methods becomeprominent
How did an early idea become a practical training approach and then face competition from another machine-learning family?

The later history was not a straight march from old ideas to modern ones. After Backpropagation helped neural networks gain attention, kernel methods became prominent classification algorithms in the 1990s and pushed neural networks back into obscurity for a time. Kernel methods are a family of classification algorithms; the best-known example is the support vector machine, or SVM. The historical contrast is useful: neural networks were developed as trainable chains of parametric operations, while kernel methods represented a competing family of classification approaches.

From Loss to a Weight Change

Training a neural network means finding weight values that produce the smallest possible loss. The gradient describes how the loss changes with respect to the collection of weights. Stochastic gradient descent uses this information to adjust the weights in the opposite direction from the gradient, because the goal is to reduce loss. The update is incremental: it modifies the current parameters rather than solving for all optimal weights in one step.

small updatereverse directionweightcurrent valueweightslightly lower valuepositive gradientloss increases in thisdirectionopposite directionloss-reducing update
How does the direction of the gradient determine the direction of a small weight change?

Reading the Direction of an Update

A training iteration produces a gradient for one weight. The gradient points toward increasing loss in the positive direction. What kind of change should stochastic gradient descent make?

Read the gradient: The gradient identifies the direction associated with how the loss changes for the current parameter.

Reverse the direction: Because the goal is to reduce loss, the update moves in the opposite direction from the gradient.

Keep the change incremental: The new weight is the previous weight after a small, gradient-guided modification, not an arbitrary replacement or an analytically solved optimum.

A gradient pointing in the positive direction produces a weight update in the opposite direction.

Why Direct Minimization Fails to Scale

In theory, a differentiable function can be minimized analytically by finding points where its derivative is zero and then checking which candidate has the lowest function value. For a neural network, this means solving the condition that the gradient of the loss with respect to the weights is zero. The difficulty is the number of unknowns: the equation has one variable for each coefficient in the network. Analytical methods might be possible for two or three variables, but real neural networks have at least a few thousand parameters and can have several tens of millions. Solving for all of them directly is therefore impractical.

becomes impractical withusesAnalyticalminimizationsolve all weight conditionstogetherMany parametersthousands to tens ofmillionsStochastic gradientdescentmake repeated small changesCurrent batchcurrent loss guides anupdate
Why does stochastic gradient descent replace one large analytical calculation with many smaller parameter changes?

Backpropagation Through Operations

Backpropagation provided a method for training chains of parametric operations through gradient-descent optimization. A neural network can be understood here as a sequence of operations whose parameters include weights. The loss at the end of the chain supplies the quantity that must be reduced, and gradient information is used to determine how the parameters should change. The important role of Backpropagation in this overview is that it makes gradient information available across the chain rather than limiting training to one isolated operation.

forward valuesforward valuesevaluate lossgradient informationback through chainInputParametric operation1weightsParametric operation2weightsLossGradient foroperation 1Gradient foroperation 2
How does training information move through successive parametric operations so that each operation can participate in an update?

The Stochastic Training Cycle

Stochastic gradient descent works with a random batch of training data rather than attempting to solve the entire optimization problem analytically. The current batch is used to evaluate the current loss. From that loss, the algorithm computes a gradient for the current parameters, then changes the parameters a little in the opposite direction from the gradient. The changed parameters are used in the next iteration. Repeating this cycle replaces one impractical global calculation with many local decisions.

evaluatecompute from lossguiderepeat with changed parametersRandom batchtraining examplesCurrent losson the batchGradientwith respect to parametersUpdated parameterssmall opposite-directionchange
What happens next as the model samples a batch, evaluates loss, computes a gradient, updates parameters, and repeats?

Following One Iteration

Trace what happens when stochastic gradient descent begins an iteration with a random batch.

Sample: The algorithm selects a random batch of training data instead of using an analytical calculation over the entire neural network optimization problem.

Evaluate: The current parameters are used with that batch to obtain the current loss value.

Compute: A gradient is computed to describe how the loss changes with respect to the parameters.

Modify: The parameters are changed a little in the opposite direction from the gradient.

Repeat: The changed parameters are carried into the next iteration, which samples another batch and performs the cycle again.

Each iteration turns a random batch and its current loss into a small, gradient-guided parameter update.

Mistakes About Gradient Descent

  • Thinking the gradient tells the algorithm which direction to update toward.

    The gradient indicates how the loss changes, while the training goal is to reduce loss.

    Fix: Update the weight in the opposite direction from the gradient.

  • Treating one update as the final solution.

    Stochastic gradient descent repeatedly makes small changes using the current batch, loss, and gradient.

    Fix: Track the repeated cycle and understand each update as one incremental step.

  • Assuming analytical minimization is practical for a real neural network.

    The number of unknowns can range from at least a few thousand parameters to several tens of millions.

    Fix: Understand stochastic gradient descent as a practical replacement for the impractical global calculation.

  • Treating kernel methods as another name for neural networks.

    Kernel methods are a family of classification algorithms, and SVM is identified as their best-known example.

    Fix: Keep the historical categories distinct: neural networks are trainable chains of parametric operations, while kernel methods are a competing classification family.

Check Your Understanding

MEDIUM

Explain the full stochastic gradient-descent cycle in your own words. Your explanation should include the random batch, the current loss, the gradient, the opposite-direction update, and the reason the cycle repeats.

Hints
  • Start with the data selected for one iteration.
  • State what the gradient describes about the loss and parameters.
  • Explain why the update uses the opposite direction.
  • Contrast repeated small updates with analytical minimization.

What do you think happens?

A parameter has a gradient pointing in one direction. Should stochastic gradient descent change that parameter in the same direction or the opposite direction?

  • The same direction
  • The opposite direction
  • It does not change the parameter
Reveal answer

Answer: The opposite direction

The gradient describes how the loss changes with respect to the parameter, and stochastic gradient descent uses the opposite direction to reduce loss.

Key Takeaways

  • Neural networks were investigated as early as the 1950s, but efficient training was needed for the approach to develop.
  • Backpropagation provided a way to train chains of parametric operations through gradient-descent optimization.
  • Analytical minimization becomes impractical because real neural networks contain thousands to tens of millions of parameters.
  • Stochastic gradient descent uses a random batch to compute a current gradient and makes a small parameter change in the opposite direction.
  • Training repeats the cycle of sampling data, evaluating loss, computing a gradient, and updating parameters.