Concepts / Gradient of a Function

Gradient of a Function

The gradient identifies the direction in which the function increases fastest around the current vector.

  • Programming

From a Current Vector to a Decision

Gradient descent does not jump directly to its final answer. It begins with an initial vector, examines the function at that current vector, and repeatedly decides which direction to move next. The gradient provides the direction of greatest local increase, so the algorithm reverses that direction to descend.

One Update in Numbers

What do you think happens?

Suppose the current vector is [10], the gradient at that vector is [2], and the positive learning-rate parameter is 0.5. The scaled movement is therefore [1]. What next vector results when the scaled movement is subtracted?

  • [11]
  • [10]
  • [9]
Reveal answer

Answer: [9]

The scaled movement is 0.5 × [2] = [1]. Reversing the gradient direction means subtracting that movement from the current vector: [10] − [1] = [9].

Applying one descent update

Current vector: [10]. Gradient at the current vector: [2]. Positive parameter η: 0.5.

Identify the direction: The gradient [2] indicates the direction in which the function increases fastest around the current vector.

Scale the direction: Multiply the gradient by η: 0.5 × [2] = [1]. The parameter controls the size of the movement.

Reverse the direction: Gradient descent subtracts the scaled gradient from the current vector: [10] − [1].

The next vector is [9].

scale by ηscalesubtract [1]subtract[10]current vector[1]scaled movement[2]gradient[9]next vector0.5η
How does the current vector change into the next vector when the gradient and learning rate are applied?

What the Gradient Tells You

For a differentiable function, the gradient identifies the direction in which the function increases fastest around the current vector.

The gradient is a direction signal. At the current vector, it points toward the greatest local increase in the function. Gradient descent needs the opposite direction because its purpose is descent rather than increase. The update therefore reverses the gradient direction and scales the movement using the positive parameter η.

gradient points towardreversemove to reduceCurrent vectorReduced functionvalueGradientfastest local increaseOpposite directiondirection used for descent
How does the gradient point toward increasing function values, and why does moving in the opposite direction provide the direction used for descent?

Reading the Update Rule

next vector = current vector − η × gradient at the current vector

combine withscalestart fromremoveproduceCurrent vectorpoint of evaluationScaled movementη × gradientNext vectorGradientdirection of fastestincreaseSubtractionreverse directionηpositive scale
How do the current vector, gradient, learning rate, and subtraction operation connect to produce the next vector?

Repeating the Update

After one update, the newly produced vector becomes the current vector for the next iteration. The algorithm evaluates the gradient again at this new point and applies the same rule again. Starting from an initial value, often written as w(1) = 0 in the description of the algorithm, the process continues through the chosen number of iterations.

evaluateapply updatebecomes currentapply update againInitial vectorw(1)Gradient at currentpointNext vectorGradient at new pointLater vector
How does one update become the starting point for the next update?

The repeated updates create a sequence of vectors. The gradient is not evaluated only once at the starting point; it is evaluated again at each newly produced vector. This is why the algorithm is an iterative process.

Choosing the Final Output

After T iterations, gradient descent needs a rule for selecting its output from the sequence of vectors. The algorithm can return the average vector, the last vector, or the best-performing vector.

Output choiceWhat it returnsSource note
Average vectorThe average of the vectors produced during the iterationsEspecially useful when gradient descent is extended to nondifferentiable functions and to the stochastic case
Last vectorThe vector produced by the final updateUses the end of the update sequence
Best-performing vectorThe vector selected as the best performer during the processSelects from the vectors produced during the iterations

Mistakes in Direction and Output

  • Moving in the same direction as the gradient

    The gradient points toward the greatest local increase, while gradient descent uses the opposite direction.

    Fix: Subtract the scaled gradient from the current vector.

  • Treating the gradient as the step size

    The gradient supplies the direction, while the positive parameter η scales the movement.

    Fix: First scale the gradient with η, then subtract the resulting movement.

  • Evaluating the gradient only at the initial vector

    Each newly produced vector becomes the current point for the next gradient evaluation.

    Fix: Re-evaluate the gradient at the current vector on every iteration.

  • Assuming the final output must be the last vector

    The algorithm can return an average vector, the last vector, or the best-performing vector.

    Fix: Choose the output rule that is part of the algorithm's definition.

Check Your Understanding

EASY

A current vector is [12]. The gradient at that vector is [4], and η is 0.25. Determine the scaled movement and the next vector. Then state which vector becomes the current vector for the next iteration.

Hints
  • Multiply η by the gradient to find the scaled movement.
  • Subtract the scaled movement from the current vector.
  • The newly produced vector becomes the current vector for the next iteration.
  1. The gradient identifies the direction in which a differentiable function increases fastest around the current vector.
  2. Gradient descent reverses the gradient direction because it is used to reduce the function value.
  3. The positive parameter η scales the movement before subtraction produces the next vector.
  4. Each next vector becomes the current vector for the following iteration, creating a sequence of vectors.
  5. After the iterations, the algorithm can return an average vector, the last vector, or the best-performing vector.

Key Takeaways

  • The gradient points toward the fastest local increase of a differentiable function.
  • Gradient descent moves opposite to that direction by subtracting a scaled gradient.
  • The current vector, gradient, positive parameter η, and subtraction operation each have a distinct role in the update.
  • The updated vector becomes the current point for the next iteration.
  • The final output may be the average vector, the last vector, or the best-performing vector.