Concepts / Step-Size Parameters in Reinforcement Learning

Step-Size Parameters in Reinforcement Learning

Sample-average updating reduces the step size according to αn(a) = 1/n.

  • Machine Learning

Why Estimates Need a Step Size

A reinforcement-learning agent updates its estimate of an action's reward behavior as new rewards arrive. The step-size parameter controls how strongly each new observation changes that estimate. This creates an important choice: should learning gradually preserve accumulated experience, or should it remain responsive to recent rewards?

The sample-average method reduces the step size as an action is sampled repeatedly. A constant step size takes the opposite approach: it continues allowing new rewards to change the estimate. Neither behavior is universally best. The appropriate choice depends partly on whether the environment is stable or changes over time.

Sample-Average Updating

For the sample-average method, the step size used after an action has been sampled n times is α_n(a) = 1/n.

more samplesmore samplesSample 1α₁(a) = 1Sample 2α₂(a) = 1/2Sample nα_n(a) = 1/n
How does α_n(a) change as the same action is sampled repeatedly?

Following the Step Size

An action is sampled repeatedly using the sample-average method. What step size is used on the first, second, and tenth samples?

First sample: Substitute n = 1 into α_n(a) = 1/n. The step size is 1.

Second sample: Substitute n = 2. The step size is 1/2.

Tenth sample: Substitute n = 10. The step size is 1/10.

The step size becomes smaller as the action is sampled more often.

This decreasing step size means that each new observation has a smaller update role as more samples have accumulated. The method is therefore designed to preserve accumulated experience rather than keep reacting equally strongly to every later reward.

Why the Rule Converges

The sample-average choice α_n(a) = 1/n satisfies both convergence conditions. The first condition requires the step sizes to have an infinite total sum. The second requires the squared step sizes to have a finite total sum.

α_n(a) = 1/n

sumsumStep sizes1 + 1/2 + 1/3 + ...Infinite sumSquared step sizes1 + 1/4 + 1/9 + ...Finite sum
Why does the sequence α_n = 1/n have an infinite sum while the sequence of squared step sizes has a finite sum?

Constant Step-Size Behavior

A constant step size keeps α at the same value instead of reducing it according to 1/n. As a result, later rewards continue to influence the estimate. The estimate does not completely settle because the second convergence condition is not met.

update with αinfluencesnext observationcontinues influenceCurrent estimateRecent rewardUpdated estimateconstant α remains activeLater reward
How does using a constant α make recent rewards continue influencing the estimate?

The continuing movement is not automatically a defect. It reflects an intentional trade-off. Sample-average updating accumulates experience and moves toward settling, while a constant step size preserves sensitivity to newer observations.

more samplesmore observationsSample averageearly estimateConstant step sizeearly estimateSample averagetends toward settlingConstant step sizecontinues adapting
What is the difference between estimates that settle as samples arrive and estimates that continue adapting?

Changing Environments

An estimate is useful only when it reflects the environment the agent currently faces. In a nonstationary environment, the conditions affecting rewards can change over time. An estimate that completely settles may preserve information that no longer describes those current conditions.

environment changesconstant α updatesEarlier conditionsinitial reward behaviorChanged conditionsreward behavior changesCurrent estimateresponds to newer rewards
How does a constant step size let an estimate follow a reward distribution that changes over time?

Consider an estimate for an action whose reward behavior changes over time. With sample-average updating, the influence of each new reward decreases as the update count grows. With a constant step size, later rewards continue influencing the estimate, so the estimate keeps varying as the environment changes.

Choosing Between the Methods

FeatureSample-average step sizeConstant step size
Step-size behaviorα_n(a) = 1/n, so it decreases with repeated samplingα remains constant
Convergence conditionsMeets both convergence conditionsDoes not meet the second convergence condition
Estimate behaviorMoves toward settling as experience accumulatesDoes not completely settle
Response to newer rewardsInfluence of each new reward decreases as the update count growsLater rewards continue to influence the estimate
Environment contextSuitable for theoretical convergence analysisDesirable when the environment is nonstationary
  • Assuming that a constant step size should eventually make the estimate settle completely.

    The ongoing response is a consequence of keeping the step size constant.

    Fix: Treat continued variation as expected behavior rather than as automatic evidence of an error.

  • Treating convergence as proof that an estimate will adapt quickly to a changing environment.

    Convergence and rapid adaptation describe different properties.

    Fix: Consider both the convergence behavior and whether the environment is nonstationary.

  • Assuming that continued movement is always undesirable.

    A completely settled estimate can preserve information that no longer matches the current environment.

    Fix: Recognize why a constant step size can be useful when reward conditions change.

Check Your Understanding

MEDIUM

Explain why α_n(a) = 1/n satisfies both convergence conditions, then describe one reason a constant step size may be preferable when reward conditions change over time.

Hints
  • Consider the sum of the step sizes and the sum of their squares.
  • Compare how the two methods treat newer rewards after many observations.
  • Relate continued variation to a nonstationary environment.

What do you think happens?

An action has been sampled many times. Which method is more likely to keep responding to a newly changing reward pattern: sample-average updating or a constant step size?

  • Sample-average updating
  • A constant step size
Reveal answer

Answer: A constant step size

With sample-average updating, the influence of each new reward decreases as the update count grows. A constant step size continues allowing later rewards to influence the estimate.

Key Takeaways

  1. The sample-average method uses α_n(a) = 1/n.
  2. This sequence has an infinite sum while its squared terms have a finite sum, so it meets both convergence conditions.
  3. A constant step size does not meet the second convergence condition, so estimates do not completely settle.
  4. Continual response to newer rewards can be useful when the environment is nonstationary.
  5. The central choice is between preserving accumulated experience and remaining responsive to changing reward conditions.

Key Takeaways

  • Sample-average updating uses α_n(a) = 1/n, causing the step size to decrease as an action is sampled repeatedly.
  • The sample-average sequence satisfies both convergence conditions because its step sizes have an infinite sum while their squares have a finite sum.
  • A constant step size keeps later rewards influential, so estimates continue changing rather than completely settling.
  • This continued adaptation is useful in nonstationary environments where current reward conditions may differ from earlier ones.