Concepts / Convergence in Reinforcement-Learning Estimates

Convergence in Reinforcement-Learning Estimates

Sample-average updating reduces the step size according to αn(a) = 1/n.

  • Machine Learning

When an Estimate Should Change

A reinforcement-learning estimate is useful only when it reflects the environment the agent currently faces. If reward conditions remain stable, an estimate can gradually settle as the agent accumulates experience. If those conditions change, however, an estimate that preserves too much old experience may no longer describe the current situation. The choice of step size controls this balance between learning from accumulated experience and responding to newer rewards.

The central question is not simply whether an estimate changes. It is whether its pattern of change is appropriate for the environment.

Visit-by-Visit Step Sizes

The sample-average method reduces the step size according to αₙ(a) = 1/n, where n is the number of times action a has been sampled. On the first visit, the step size is 1. On the second visit, it is 1/2. On the third visit, it is 1/3. As the visit count grows, each new reward receives a smaller step size.

next visitnext visitcontinued visitsVisit 1α₁(a) = 1Visit 2α₂(a) = 1/2Visit 3α₃(a) = 1/3Visit nαₙ(a) = 1/n
What happens to the update step size as the same action is sampled more times?

Following One Action's Step Size

An action has been sampled repeatedly using the sample-average method. Determine the step size on the first three visits and on visit n.

First visit: Substitute n = 1 into αₙ(a) = 1/n. The step size is α₁(a) = 1.

Second visit: Substitute n = 2. The step size is α₂(a) = 1/2.

Third visit: Substitute n = 3. The step size is α₃(a) = 1/3.

General visit: On visit n, the step size is 1/n, so it becomes smaller as the number of visits increases.

The sample-average rule gives progressively smaller step sizes: 1, 1/2, 1/3, and so on.

Why Sample Averages Meet Both Conditions

The sample-average choice αₙ(a) = 1/n meets both convergence conditions. The first condition requires the total of the step sizes to remain continually accumulated rather than ending at a finite total. The second requires the total of the squared step sizes to be finite. For the sample-average sequence, the step sizes are 1/n, while the squared step sizes are 1/n². Thus the ordinary step sizes continue to contribute across visits, but their squared contributions become sufficiently small for the squared total to remain finite.

αₙ(a) = 1/n

sumsumStep sizes1, 1/2, 1/3, ...Ongoing totaldoes not end at a finitetotalSquared steps1, 1/4, 1/9, ...Finite totalsquared contributionssettle
Why does αₙ(a) = 1/n have an ongoing total of step sizes but a finite total of squared step sizes?

Constant Step Sizes

A constant step size uses the same α on every update instead of shrinking according to 1/n. Because the step size does not become smaller with experience, later rewards continue to influence the estimate. The estimate therefore does not completely settle in the same way as the sample-average case, because a constant step size does not meet the second convergence condition.

time passesnew observationsconstant α keeps influenceEarly rewardsinitial environmentReward conditionschangeenvironment becomesnonstationaryRecent rewardsnew environmentChanging estimatecontinues responding
How does a constant step size allow newer rewards to keep changing an estimate after earlier observations have become less representative?

Stable Versus Continually Changing Estimates

Compare an action estimate updated with αₙ(a) = 1/n against one updated with a constant step size when the action's reward behavior changes over time.

Sample-average update: Each new reward receives a smaller step size as the action is sampled more times. The influence of later rewards decreases with the visit count.

Constant-step-size update: Every update uses the same α. Later rewards continue to influence the estimate rather than becoming nearly irrelevant solely because many earlier visits occurred.

Changing environment: If reward behavior changes, the constant-step-size estimate can continue responding to the newer observations, while the sample-average method remains centered on accumulated experience.

The sample-average method is suited to convergence analysis, while a constant step size preserves ongoing responsiveness to newer rewards.

Convergence and Responsiveness

Step-size choiceBehavior over visitsUse in changing environments
Sample average: αₙ(a) = 1/nStep size shrinks as visits accumulate; the method meets both convergence conditions.Older experience continues to be accumulated, so adaptation to newly changed conditions may be slow.
Constant αThe same step size is used repeatedly; the estimate does not completely settle.Later rewards continue to influence the estimate, which is useful when reward conditions change.

These methods represent different priorities. Sample-average updating is designed around accumulating experience and satisfies both convergence conditions. A constant step size gives up that same convergence behavior so that the estimate can remain sensitive to newer observations. In a nonstationary environment, continued movement is not automatically a defect: an estimate that completely settles may preserve information that no longer describes current conditions.

Common Reasoning Errors

  • Assuming that every useful estimate should eventually stop changing.

    A nonstationary environment can make older information less representative of current reward conditions.

    Fix: Distinguish between convergence behavior and responsiveness to newer rewards.

  • Confusing a constant step size with the sample-average rule.

    The sample-average step size shrinks with the visit count, whereas a constant step size remains the same on every update.

    Fix: Check whether the update uses αₙ(a) = 1/n or one fixed α.

  • Assuming that satisfying the convergence conditions guarantees rapid learning.

    A sequence can satisfy the conditions and still converge slowly or require tuning.

    Fix: Treat convergence guarantees and convergence rate as separate considerations.

Check Your Understanding

MEDIUM

An action has been sampled many times. You are deciding between αₙ(a) = 1/n and a constant α because the environment may change. Explain what happens to the influence of new rewards under each choice, identify which choice meets both convergence conditions, and state why the other choice may be preferable in a nonstationary environment.

Hints
  • Track how the step size behaves as n increases.
  • Connect the sample-average sequence to the two convergence conditions.
  • Consider whether old experience must always remain representative of current conditions.

Key Takeaways

  1. The sample-average method uses αₙ(a) = 1/n, so the step size shrinks as an action is sampled more often.
  2. This sequence meets both convergence conditions: the total of the step sizes remains ongoing, while the total of their squares is finite.
  3. A constant step size does not meet the second convergence condition, so its estimate does not completely settle.
  4. Continued variation can be useful in a nonstationary environment because newer rewards may better reflect the conditions the agent currently faces.
  5. Convergence and responsiveness are different goals; the appropriate step-size choice depends on whether accumulated experience or adaptation to change is more important.

Key Takeaways

  • Sample-average updating uses αₙ(a) = 1/n.
  • The sample-average sequence satisfies both convergence conditions, although convergence may be slow.
  • A constant step size keeps later rewards influential and therefore prevents complete settling.
  • That continuing variation is useful when the reward environment is nonstationary.
  • Step-size selection balances theoretical convergence against responsiveness to current conditions.