Convergence in Optimization
Variable step size SGD changes the update scale during optimization.
Why Movement Changes
In stochastic gradient descent, the step size controls how far an update moves during optimization. Using one constant step size gives every update the same scale. A variable step size changes that scale according to the iteration number: optimization can take larger steps earlier and more careful steps later.
The central idea is a schedule for the step size. Instead of selecting one value and reusing it forever, SGD computes the step size for the current iteration. The schedule η_t = B / √t makes the step size decrease as t increases. This changes the character of the optimization process: early updates can be larger, while later updates are smaller and more careful near the minimum.
Reading the Schedule
η_t = B / √t
The iteration number appears in the denominator. At an early iteration, √t is relatively small, so the resulting step size is relatively large. At a later iteration, √t is larger, so the resulting step size is smaller. The schedule therefore encodes a deliberate change from larger movement to more careful movement.
A Numerical Trace
Evaluating η_t at Several Iterations
Suppose the schedule uses B = 4. Find the step size at iterations 1, 4, 9, and 16.
Iteration 1: η₁ = 4 / √1 = 4. The schedule gives a relatively large step size at this early iteration.
Iteration 4: η₄ = 4 / √4 = 4 / 2 = 2. The step size is smaller than it was at iteration 1.
Iteration 9: η₉ = 4 / √9 = 4 / 3. The step size has decreased again.
Iteration 16: η₁₆ = 4 / √16 = 4 / 4 = 1. The later step is smaller than all the earlier steps in this example.
As t increases from 1 to 16, the step size decreases from 4 to 1. The schedule changes the update scale rather than keeping it constant.
| Iteration t | √t | Step size η_t when B = 4 | Optimization phase |
|---|---|---|---|
| 1 | 1 | 4 | Early |
| 4 | 2 | 2 | Early |
| 9 | 3 | 4 / 3 | Later |
| 16 | 4 | 1 | Later |
Approaching the Minimum
A constant step size applies the same movement scale to every update. That can make later movement less careful because the update does not become smaller as the iteration number grows. With η_t = B / √t, the update scale decreases over time. The source describes this as supporting more careful movement near the minimum and helping avoid overshooting.
Update Scale Over Time
The schedule must be evaluated using the current iteration. At each step, t identifies the current position in the optimization process, and η_t supplies the scale for that update. Recomputing the schedule is what makes the method variable-step SGD.
Common Scheduling Mistakes
Describing η_t = B / √t as a constant step size.
The schedule explicitly depends on t, and it decreases as t increases.
Fix:
Compute η_t from the current iteration number each time the update scale is needed.Saying that later iterations use larger steps.
Increasing t increases √t in the denominator, so the quotient becomes smaller.
Fix:
Describe the schedule as larger earlier steps followed by smaller later steps.Explaining only that the numerical value changes.
The purpose of the schedule is to support more careful movement near the minimum and help avoid overshooting.
Fix:
Connect the smaller later step size to more careful movement near the minimum.
When explaining or applying this schedule, always identify the current iteration t, evaluate η_t = B / √t, and then interpret the result as the scale of the current SGD update. This keeps the mathematical schedule and its optimization meaning connected.
Check Your Understanding
Let B remain fixed. Compare the step sizes at t = 1 and t = 100 using η_t = B / √t. Which iteration uses the larger step size, and why does this behavior support more careful movement later in optimization?
Hints
- Compare the denominators √1 and √100.
- Remember that a larger denominator makes the quotient smaller.
- Relate the smaller later value to movement near the minimum.
What do you think happens?
If B stays fixed, which step size is larger: η₄ or η₁₆?
Reveal answer
Answer: η₄
η₄ = B / 2, while η₁₆ = B / 4. The larger denominator at t = 16 produces the smaller step size.
Key Takeaways
- SGD uses its step size to control how far each update moves.
- The schedule η_t = B / √t decreases the step size as the iteration number increases.
- The schedule allows larger movement earlier and more careful movement later.
- Smaller later steps can help reduce overshooting near the minimum.
- The step size must be recomputed from the current iteration rather than reused as one constant value.
Key Takeaways
- Variable-step SGD changes the update scale during optimization.
- For η_t = B / √t, increasing t makes η_t smaller.
- The schedule gives larger steps early and more careful steps later.
- Decreasing the step size can help avoid overshooting near a minimum.
- The current iteration must be used to recompute the step size.