Gradient of a Function
The gradient identifies the direction in which the function increases fastest around the current vector.
From a Current Vector to a Decision
Gradient descent does not jump directly to its final answer. It begins with an initial vector, examines the function at that current vector, and repeatedly decides which direction to move next. The gradient provides the direction of greatest local increase, so the algorithm reverses that direction to descend.
One Update in Numbers
What do you think happens?
Suppose the current vector is [10], the gradient at that vector is [2], and the positive learning-rate parameter is 0.5. The scaled movement is therefore [1]. What next vector results when the scaled movement is subtracted?
Reveal answer
Answer: [9]
The scaled movement is 0.5 × [2] = [1]. Reversing the gradient direction means subtracting that movement from the current vector: [10] − [1] = [9].
Applying one descent update
Current vector: [10]. Gradient at the current vector: [2]. Positive parameter η: 0.5.
Identify the direction: The gradient [2] indicates the direction in which the function increases fastest around the current vector.
Scale the direction: Multiply the gradient by η: 0.5 × [2] = [1]. The parameter controls the size of the movement.
Reverse the direction: Gradient descent subtracts the scaled gradient from the current vector: [10] − [1].
The next vector is [9].
What the Gradient Tells You
For a differentiable function, the gradient identifies the direction in which the function increases fastest around the current vector.
The gradient is a direction signal. At the current vector, it points toward the greatest local increase in the function. Gradient descent needs the opposite direction because its purpose is descent rather than increase. The update therefore reverses the gradient direction and scales the movement using the positive parameter η.
Reading the Update Rule
next vector = current vector − η × gradient at the current vector
Repeating the Update
After one update, the newly produced vector becomes the current vector for the next iteration. The algorithm evaluates the gradient again at this new point and applies the same rule again. Starting from an initial value, often written as w(1) = 0 in the description of the algorithm, the process continues through the chosen number of iterations.
The repeated updates create a sequence of vectors. The gradient is not evaluated only once at the starting point; it is evaluated again at each newly produced vector. This is why the algorithm is an iterative process.
Choosing the Final Output
After T iterations, gradient descent needs a rule for selecting its output from the sequence of vectors. The algorithm can return the average vector, the last vector, or the best-performing vector.
| Output choice | What it returns | Source note |
|---|---|---|
| Average vector | The average of the vectors produced during the iterations | Especially useful when gradient descent is extended to nondifferentiable functions and to the stochastic case |
| Last vector | The vector produced by the final update | Uses the end of the update sequence |
| Best-performing vector | The vector selected as the best performer during the process | Selects from the vectors produced during the iterations |
Mistakes in Direction and Output
Moving in the same direction as the gradient
The gradient points toward the greatest local increase, while gradient descent uses the opposite direction.
Fix:
Subtract the scaled gradient from the current vector.Treating the gradient as the step size
The gradient supplies the direction, while the positive parameter η scales the movement.
Fix:
First scale the gradient with η, then subtract the resulting movement.Evaluating the gradient only at the initial vector
Each newly produced vector becomes the current point for the next gradient evaluation.
Fix:
Re-evaluate the gradient at the current vector on every iteration.Assuming the final output must be the last vector
The algorithm can return an average vector, the last vector, or the best-performing vector.
Fix:
Choose the output rule that is part of the algorithm's definition.
Check Your Understanding
A current vector is [12]. The gradient at that vector is [4], and η is 0.25. Determine the scaled movement and the next vector. Then state which vector becomes the current vector for the next iteration.
Hints
- Multiply η by the gradient to find the scaled movement.
- Subtract the scaled movement from the current vector.
- The newly produced vector becomes the current vector for the next iteration.
- The gradient identifies the direction in which a differentiable function increases fastest around the current vector.
- Gradient descent reverses the gradient direction because it is used to reduce the function value.
- The positive parameter η scales the movement before subtraction produces the next vector.
- Each next vector becomes the current vector for the following iteration, creating a sequence of vectors.
- After the iterations, the algorithm can return an average vector, the last vector, or the best-performing vector.
Key Takeaways
- The gradient points toward the fastest local increase of a differentiable function.
- Gradient descent moves opposite to that direction by subtracting a scaled gradient.
- The current vector, gradient, positive parameter η, and subtraction operation each have a distinct role in the update.
- The updated vector becomes the current point for the next iteration.
- The final output may be the average vector, the last vector, or the best-performing vector.