Convergence in Reinforcement-Learning Estimates
Sample-average updating reduces the step size according to αn(a) = 1/n.
When an Estimate Should Change
A reinforcement-learning estimate is useful only when it reflects the environment the agent currently faces. If reward conditions remain stable, an estimate can gradually settle as the agent accumulates experience. If those conditions change, however, an estimate that preserves too much old experience may no longer describe the current situation. The choice of step size controls this balance between learning from accumulated experience and responding to newer rewards.
The central question is not simply whether an estimate changes. It is whether its pattern of change is appropriate for the environment.
Visit-by-Visit Step Sizes
The sample-average method reduces the step size according to αₙ(a) = 1/n, where n is the number of times action a has been sampled. On the first visit, the step size is 1. On the second visit, it is 1/2. On the third visit, it is 1/3. As the visit count grows, each new reward receives a smaller step size.
Following One Action's Step Size
An action has been sampled repeatedly using the sample-average method. Determine the step size on the first three visits and on visit n.
First visit: Substitute n = 1 into αₙ(a) = 1/n. The step size is α₁(a) = 1.
Second visit: Substitute n = 2. The step size is α₂(a) = 1/2.
Third visit: Substitute n = 3. The step size is α₃(a) = 1/3.
General visit: On visit n, the step size is 1/n, so it becomes smaller as the number of visits increases.
The sample-average rule gives progressively smaller step sizes: 1, 1/2, 1/3, and so on.
Why Sample Averages Meet Both Conditions
The sample-average choice αₙ(a) = 1/n meets both convergence conditions. The first condition requires the total of the step sizes to remain continually accumulated rather than ending at a finite total. The second requires the total of the squared step sizes to be finite. For the sample-average sequence, the step sizes are 1/n, while the squared step sizes are 1/n². Thus the ordinary step sizes continue to contribute across visits, but their squared contributions become sufficiently small for the squared total to remain finite.
αₙ(a) = 1/n
Constant Step Sizes
A constant step size uses the same α on every update instead of shrinking according to 1/n. Because the step size does not become smaller with experience, later rewards continue to influence the estimate. The estimate therefore does not completely settle in the same way as the sample-average case, because a constant step size does not meet the second convergence condition.
Stable Versus Continually Changing Estimates
Compare an action estimate updated with αₙ(a) = 1/n against one updated with a constant step size when the action's reward behavior changes over time.
Sample-average update: Each new reward receives a smaller step size as the action is sampled more times. The influence of later rewards decreases with the visit count.
Constant-step-size update: Every update uses the same α. Later rewards continue to influence the estimate rather than becoming nearly irrelevant solely because many earlier visits occurred.
Changing environment: If reward behavior changes, the constant-step-size estimate can continue responding to the newer observations, while the sample-average method remains centered on accumulated experience.
The sample-average method is suited to convergence analysis, while a constant step size preserves ongoing responsiveness to newer rewards.
Convergence and Responsiveness
| Step-size choice | Behavior over visits | Use in changing environments |
|---|---|---|
| Sample average: αₙ(a) = 1/n | Step size shrinks as visits accumulate; the method meets both convergence conditions. | Older experience continues to be accumulated, so adaptation to newly changed conditions may be slow. |
| Constant α | The same step size is used repeatedly; the estimate does not completely settle. | Later rewards continue to influence the estimate, which is useful when reward conditions change. |
These methods represent different priorities. Sample-average updating is designed around accumulating experience and satisfies both convergence conditions. A constant step size gives up that same convergence behavior so that the estimate can remain sensitive to newer observations. In a nonstationary environment, continued movement is not automatically a defect: an estimate that completely settles may preserve information that no longer describes current conditions.
Common Reasoning Errors
Assuming that every useful estimate should eventually stop changing.
A nonstationary environment can make older information less representative of current reward conditions.
Fix:
Distinguish between convergence behavior and responsiveness to newer rewards.Confusing a constant step size with the sample-average rule.
The sample-average step size shrinks with the visit count, whereas a constant step size remains the same on every update.
Fix:
Check whether the update uses αₙ(a) = 1/n or one fixed α.Assuming that satisfying the convergence conditions guarantees rapid learning.
A sequence can satisfy the conditions and still converge slowly or require tuning.
Fix:
Treat convergence guarantees and convergence rate as separate considerations.
Check Your Understanding
An action has been sampled many times. You are deciding between αₙ(a) = 1/n and a constant α because the environment may change. Explain what happens to the influence of new rewards under each choice, identify which choice meets both convergence conditions, and state why the other choice may be preferable in a nonstationary environment.
Hints
- Track how the step size behaves as n increases.
- Connect the sample-average sequence to the two convergence conditions.
- Consider whether old experience must always remain representative of current conditions.
Key Takeaways
- The sample-average method uses αₙ(a) = 1/n, so the step size shrinks as an action is sampled more often.
- This sequence meets both convergence conditions: the total of the step sizes remains ongoing, while the total of their squares is finite.
- A constant step size does not meet the second convergence condition, so its estimate does not completely settle.
- Continued variation can be useful in a nonstationary environment because newer rewards may better reflect the conditions the agent currently faces.
- Convergence and responsiveness are different goals; the appropriate step-size choice depends on whether accumulated experience or adaptation to change is more important.
Key Takeaways
- Sample-average updating uses αₙ(a) = 1/n.
- The sample-average sequence satisfies both convergence conditions, although convergence may be slow.
- A constant step size keeps later rewards influential and therefore prevents complete settling.
- That continuing variation is useful when the reward environment is nonstationary.
- Step-size selection balances theoretical convergence against responsiveness to current conditions.