Concepts / Neural Network Training

Neural Network Training

Analytical minimization identifies derivative-zero points, but it is not practical for large neural networks.

  • Programming

From Optimal Weights to Incremental Learning

Training a neural network means finding weight values that produce the smallest possible loss. One tempting approach is to solve for the best weights directly. For a differentiable function, analytical minimization looks for points where the derivative is zero and then checks which candidate has the lowest function value. For a neural network, this would mean solving a gradient-equals-zero condition for all of the network's weights at once.

Why Direct Minimization Does Not Scale

The obstacle is the number of unknowns. The equation has one variable for each coefficient in the network. Analytical methods might be possible for a problem with two or three variables, but real neural networks have at least a few thousand parameters and can have several tens of millions. Solving the complete gradient-equals-zero system would therefore require a global calculation over an enormous number of variables.

ApproachWhat it attemptsHow parameters change
Analytical minimizationFind derivative-zero points and compare their function valuesAttempts to solve the optimization problem globally
Stochastic gradient descentUse the current loss and gradient from a random batchMakes many small changes over repeated iterations

A Batch Becomes an Update Signal

Stochastic gradient descent does not attempt to solve the entire neural-network optimization problem analytically. Instead, it selects a random batch of training data. The network uses that current batch to evaluate the current loss, compute a gradient, and guide a small parameter update. The changed parameters are then used in the next iteration.

used by networkevaluateddifferentiatedguides updateRandom batchtraining dataPredictionsCurrent lossGradientUpdated parameters
How does a randomly selected batch move through prediction, loss evaluation, gradient computation, and parameter updating?

Following One Batch

Trace what happens during one stochastic gradient descent iteration after a random batch has been selected.

Select: A random batch of training data is chosen instead of using an analytical calculation over the entire optimization problem.

Evaluate: The current network parameters are used with that batch to determine the current loss value.

Compute: Because the loss-related function is differentiable, a gradient can be computed for the parameters.

Modify: The parameters are changed a small amount in the direction opposite to the gradient.

Continue: The modified parameters become the parameters used in the next iteration.

One batch does not finish training. It supplies the current loss and gradient for one incremental update.

Reading the Gradient Direction

For a collection of weights, the gradient describes how the loss changes with respect to the parameters. This makes the gradient the link between the current loss and the next parameter update. Stochastic gradient descent uses the opposite direction from the gradient because the goal is to reduce loss, not move in the direction associated with increasing it.

Suppose one parameter has a gradient pointing toward increasing loss. The update moves that parameter in the opposite direction. The new value is therefore not an arbitrary replacement: it is the old value after a small, gradient-guided modification. The same directional reasoning applies across the collection of weights.

modified using opposite directionguides directionWeight wcurrent valueWeight wsmall opposite-gradientchangeGradientdirection from loss
How does the gradient determine the direction and amount of a weight change to reduce the loss?

The Repeating Training Cycle

The training algorithm repeats the same reasoning. It samples a random batch, evaluates the loss for the current parameters on that batch, computes the gradient, changes the parameters in the opposite direction, and then starts the next iteration with those changed parameters. Each cycle uses the current state rather than solving for all optimal weights in one step.

current batchdifferentiateguide updateuse changed parametersrepeatSample batchEvaluate lossCompute gradientModify parametersopposite gradient directionNext iteration
What happens next as training repeatedly samples a batch, evaluates loss, computes a gradient, and modifies the network parameters?

A single iteration should not be confused with a complete solution. Stochastic gradient descent replaces one impractical global calculation with many small parameter changes, so the important object to track is the sequence of updated parameter states.

Mistakes About Gradient Descent

  • Treating the gradient as the new weight value.

    The gradient describes how the loss changes with respect to the parameter. It is the link between the current loss and the update, not an arbitrary replacement value for the parameter.

    Fix: Keep the current parameter and modify it a small amount in the opposite direction from the gradient.

  • Assuming the update follows the gradient direction.

    The goal is to reduce loss.

    Fix: Use the direction opposite to the gradient.

  • Assuming one random batch solves the whole training problem.

    A batch guides one incremental update, and training repeats the cycle with changed parameters.

    Fix: View each batch as one step in a repeated process.

  • Expecting analytical minimization to remain practical as the network grows.

    The global equation has one variable for each network coefficient, making the calculation impractical for real neural networks.

    Fix: Understand stochastic gradient descent as a replacement for that global calculation with many small parameter changes.

Check the Update Logic

MEDIUM

Explain the complete path from a randomly selected batch to the next parameter state. In your explanation, identify where the current loss is evaluated, what the gradient describes, why the update uses the opposite direction, and why the process must repeat.

Hints
  • Start with the batch rather than with the full training set.
  • Connect the current loss to the computed gradient.
  • Describe the parameter after the update as a modified version of its previous value.

What do you think happens?

A training iteration has computed a gradient for a parameter. Should stochastic gradient descent change the parameter in the same direction as the gradient or the opposite direction?

  • The same direction
  • The opposite direction
Reveal answer

Answer: The opposite direction

The gradient provides the direction associated with change in the loss, while the training goal is to reduce loss.

Training in One Sentence

  1. Analytical minimization looks for derivative-zero points, but the number of parameters in a real neural network makes the global calculation impractical.
  2. Stochastic gradient descent uses a random batch to obtain the current loss and guide one incremental update.
  3. The gradient describes how the loss changes with respect to the weights.
  4. Weights are modified in the opposite direction from the gradient to reduce loss.
  5. Training repeats batch sampling, loss evaluation, gradient computation, and parameter modification.

Key Takeaways

  • Direct analytical minimization is impractical because a neural network can contain thousands to tens of millions of parameters.
  • A random batch supplies the current data used to evaluate loss and compute a gradient.
  • The gradient guides a small parameter change, and the change uses the opposite gradient direction.
  • Training is a repeated cycle in which updated parameters become the starting point for the next iteration.