Step-Size Parameters in Reinforcement Learning
Sample-average updating reduces the step size according to αn(a) = 1/n.
Why Estimates Need a Step Size
A reinforcement-learning agent updates its estimate of an action's reward behavior as new rewards arrive. The step-size parameter controls how strongly each new observation changes that estimate. This creates an important choice: should learning gradually preserve accumulated experience, or should it remain responsive to recent rewards?
The sample-average method reduces the step size as an action is sampled repeatedly. A constant step size takes the opposite approach: it continues allowing new rewards to change the estimate. Neither behavior is universally best. The appropriate choice depends partly on whether the environment is stable or changes over time.
Sample-Average Updating
For the sample-average method, the step size used after an action has been sampled n times is α_n(a) = 1/n.
Following the Step Size
An action is sampled repeatedly using the sample-average method. What step size is used on the first, second, and tenth samples?
First sample: Substitute n = 1 into α_n(a) = 1/n. The step size is 1.
Second sample: Substitute n = 2. The step size is 1/2.
Tenth sample: Substitute n = 10. The step size is 1/10.
The step size becomes smaller as the action is sampled more often.
This decreasing step size means that each new observation has a smaller update role as more samples have accumulated. The method is therefore designed to preserve accumulated experience rather than keep reacting equally strongly to every later reward.
Why the Rule Converges
The sample-average choice α_n(a) = 1/n satisfies both convergence conditions. The first condition requires the step sizes to have an infinite total sum. The second requires the squared step sizes to have a finite total sum.
α_n(a) = 1/n
Constant Step-Size Behavior
A constant step size keeps α at the same value instead of reducing it according to 1/n. As a result, later rewards continue to influence the estimate. The estimate does not completely settle because the second convergence condition is not met.
The continuing movement is not automatically a defect. It reflects an intentional trade-off. Sample-average updating accumulates experience and moves toward settling, while a constant step size preserves sensitivity to newer observations.
Changing Environments
An estimate is useful only when it reflects the environment the agent currently faces. In a nonstationary environment, the conditions affecting rewards can change over time. An estimate that completely settles may preserve information that no longer describes those current conditions.
Consider an estimate for an action whose reward behavior changes over time. With sample-average updating, the influence of each new reward decreases as the update count grows. With a constant step size, later rewards continue influencing the estimate, so the estimate keeps varying as the environment changes.
Choosing Between the Methods
| Feature | Sample-average step size | Constant step size |
|---|---|---|
| Step-size behavior | α_n(a) = 1/n, so it decreases with repeated sampling | α remains constant |
| Convergence conditions | Meets both convergence conditions | Does not meet the second convergence condition |
| Estimate behavior | Moves toward settling as experience accumulates | Does not completely settle |
| Response to newer rewards | Influence of each new reward decreases as the update count grows | Later rewards continue to influence the estimate |
| Environment context | Suitable for theoretical convergence analysis | Desirable when the environment is nonstationary |
Assuming that a constant step size should eventually make the estimate settle completely.
The ongoing response is a consequence of keeping the step size constant.
Fix:
Treat continued variation as expected behavior rather than as automatic evidence of an error.Treating convergence as proof that an estimate will adapt quickly to a changing environment.
Convergence and rapid adaptation describe different properties.
Fix:
Consider both the convergence behavior and whether the environment is nonstationary.Assuming that continued movement is always undesirable.
A completely settled estimate can preserve information that no longer matches the current environment.
Fix:
Recognize why a constant step size can be useful when reward conditions change.
Check Your Understanding
Explain why α_n(a) = 1/n satisfies both convergence conditions, then describe one reason a constant step size may be preferable when reward conditions change over time.
Hints
- Consider the sum of the step sizes and the sum of their squares.
- Compare how the two methods treat newer rewards after many observations.
- Relate continued variation to a nonstationary environment.
What do you think happens?
An action has been sampled many times. Which method is more likely to keep responding to a newly changing reward pattern: sample-average updating or a constant step size?
Reveal answer
Answer: A constant step size
With sample-average updating, the influence of each new reward decreases as the update count grows. A constant step size continues allowing later rewards to influence the estimate.
Key Takeaways
- The sample-average method uses α_n(a) = 1/n.
- This sequence has an infinite sum while its squared terms have a finite sum, so it meets both convergence conditions.
- A constant step size does not meet the second convergence condition, so estimates do not completely settle.
- Continual response to newer rewards can be useful when the environment is nonstationary.
- The central choice is between preserving accumulated experience and remaining responsive to changing reward conditions.
Key Takeaways
- Sample-average updating uses α_n(a) = 1/n, causing the step size to decrease as an action is sampled repeatedly.
- The sample-average sequence satisfies both convergence conditions because its step sizes have an infinite sum while their squares have a finite sum.
- A constant step size keeps later rewards influential, so estimates continue changing rather than completely settling.
- This continued adaptation is useful in nonstationary environments where current reward conditions may differ from earlier ones.