Gradient Descent
Early neural networks were explored decades before modern deep learning, but their development was constrained by training difficulty.
Try it: Gradient Descent
How gradient descent minimises a loss by repeatedly stepping against the gradient, and how the learning rate controls whether it converges, oscillates or diverges.
How it works
- The loss is f(x) = (x − 3)² + 1, lowest at x = 3.
- At the current x, compute the gradient f′(x) = 2(x − 3).
- Update x ← x − learning rate × gradient.
- Repeat. Small learning rates creep in, good ones converge quickly, rates above 1 overshoot further each step and diverge.
Default run (40 steps): Start at x = -3. Loss f(x) = (x − 3)² + 1 = 37. Learning rate 0.1. … Update 39: gradient f′ = 2(x − 3) = -0.0025; step = −0.1 × -0.0025 = 0.0002; new x = 2.999; loss 1 → 1. x is within 0.001 of the minimum at x = 3 — converged.
Simplified: One parameter and a perfectly smooth loss. Real models optimise millions of parameters on noisy losses.
Loading the simulation…
Why Training Changed the Story
Neural networks were investigated in simple forms as early as the 1950s, decades before modern deep learning. The central idea existed, but the approach faced a practical obstacle: researchers did not yet have an efficient way to train large neural networks. In the mid-1980s, multiple people independently rediscovered Backpropagation. Combined with gradient-descent optimization, it provided a way to train chains of parametric operations. This changed the practical importance of neural networks because the challenge was no longer only imagining a network; it was finding useful values for its weights.
The later history was not a straight march from old ideas to modern ones. After Backpropagation helped neural networks gain attention, kernel methods became prominent classification algorithms in the 1990s and pushed neural networks back into obscurity for a time. Kernel methods are a family of classification algorithms; the best-known example is the support vector machine, or SVM. The historical contrast is useful: neural networks were developed as trainable chains of parametric operations, while kernel methods represented a competing family of classification approaches.
From Loss to a Weight Change
Training a neural network means finding weight values that produce the smallest possible loss. The gradient describes how the loss changes with respect to the collection of weights. Stochastic gradient descent uses this information to adjust the weights in the opposite direction from the gradient, because the goal is to reduce loss. The update is incremental: it modifies the current parameters rather than solving for all optimal weights in one step.
Reading the Direction of an Update
A training iteration produces a gradient for one weight. The gradient points toward increasing loss in the positive direction. What kind of change should stochastic gradient descent make?
Read the gradient: The gradient identifies the direction associated with how the loss changes for the current parameter.
Reverse the direction: Because the goal is to reduce loss, the update moves in the opposite direction from the gradient.
Keep the change incremental: The new weight is the previous weight after a small, gradient-guided modification, not an arbitrary replacement or an analytically solved optimum.
A gradient pointing in the positive direction produces a weight update in the opposite direction.
Why Direct Minimization Fails to Scale
In theory, a differentiable function can be minimized analytically by finding points where its derivative is zero and then checking which candidate has the lowest function value. For a neural network, this means solving the condition that the gradient of the loss with respect to the weights is zero. The difficulty is the number of unknowns: the equation has one variable for each coefficient in the network. Analytical methods might be possible for two or three variables, but real neural networks have at least a few thousand parameters and can have several tens of millions. Solving for all of them directly is therefore impractical.
Backpropagation Through Operations
Backpropagation provided a method for training chains of parametric operations through gradient-descent optimization. A neural network can be understood here as a sequence of operations whose parameters include weights. The loss at the end of the chain supplies the quantity that must be reduced, and gradient information is used to determine how the parameters should change. The important role of Backpropagation in this overview is that it makes gradient information available across the chain rather than limiting training to one isolated operation.
The Stochastic Training Cycle
Stochastic gradient descent works with a random batch of training data rather than attempting to solve the entire optimization problem analytically. The current batch is used to evaluate the current loss. From that loss, the algorithm computes a gradient for the current parameters, then changes the parameters a little in the opposite direction from the gradient. The changed parameters are used in the next iteration. Repeating this cycle replaces one impractical global calculation with many local decisions.
Following One Iteration
Trace what happens when stochastic gradient descent begins an iteration with a random batch.
Sample: The algorithm selects a random batch of training data instead of using an analytical calculation over the entire neural network optimization problem.
Evaluate: The current parameters are used with that batch to obtain the current loss value.
Compute: A gradient is computed to describe how the loss changes with respect to the parameters.
Modify: The parameters are changed a little in the opposite direction from the gradient.
Repeat: The changed parameters are carried into the next iteration, which samples another batch and performs the cycle again.
Each iteration turns a random batch and its current loss into a small, gradient-guided parameter update.
Mistakes About Gradient Descent
Thinking the gradient tells the algorithm which direction to update toward.
The gradient indicates how the loss changes, while the training goal is to reduce loss.
Fix:
Update the weight in the opposite direction from the gradient.Treating one update as the final solution.
Stochastic gradient descent repeatedly makes small changes using the current batch, loss, and gradient.
Fix:
Track the repeated cycle and understand each update as one incremental step.Assuming analytical minimization is practical for a real neural network.
The number of unknowns can range from at least a few thousand parameters to several tens of millions.
Fix:
Understand stochastic gradient descent as a practical replacement for the impractical global calculation.Treating kernel methods as another name for neural networks.
Kernel methods are a family of classification algorithms, and SVM is identified as their best-known example.
Fix:
Keep the historical categories distinct: neural networks are trainable chains of parametric operations, while kernel methods are a competing classification family.
Check Your Understanding
Explain the full stochastic gradient-descent cycle in your own words. Your explanation should include the random batch, the current loss, the gradient, the opposite-direction update, and the reason the cycle repeats.
Hints
- Start with the data selected for one iteration.
- State what the gradient describes about the loss and parameters.
- Explain why the update uses the opposite direction.
- Contrast repeated small updates with analytical minimization.
What do you think happens?
A parameter has a gradient pointing in one direction. Should stochastic gradient descent change that parameter in the same direction or the opposite direction?
Reveal answer
Answer: The opposite direction
The gradient describes how the loss changes with respect to the parameter, and stochastic gradient descent uses the opposite direction to reduce loss.
Key Takeaways
- Neural networks were investigated as early as the 1950s, but efficient training was needed for the approach to develop.
- Backpropagation provided a way to train chains of parametric operations through gradient-descent optimization.
- Analytical minimization becomes impractical because real neural networks contain thousands to tens of millions of parameters.
- Stochastic gradient descent uses a random batch to compute a current gradient and makes a small parameter change in the opposite direction.
- Training repeats the cycle of sampling data, evaluating loss, computing a gradient, and updating parameters.