Concepts / Nonstationary Problems and Weighted Averages

Nonstationary Problems and Weighted Averages

The method gives greater weight to recent rewards and smaller weight to older rewards.

  • Machine Learning

When the Past Becomes Less Useful

In a nonstationary problem, the reward situation can change over time. Because of that, an old reward may not be as useful for estimating the present as a recent reward. An exponential recency-weighted average addresses this by giving recent rewards more influence and older rewards progressively less influence.

Recent Rewards Lead the Estimate

The method assigns a weight to every observed reward. The most recent reward has the greatest influence among the observed rewards. A reward one step farther in the past receives one additional factor of 1 − α. A reward another step farther back receives another factor of 1 − α. Since 1 − α is less than 1 in this weighting pattern, repeated multiplication makes the weights smaller.

one step olderone step olderone step olderRₙαRₙ₊₁αRₙ₋₁α(1 − α)Rₙα(1 − α)Rₙ₋₁α(1 − α)²
How do the weights assigned to recent and older rewards differ as another reward is observed?

The diagram shows a pattern rather than a particular numerical calculation. When a new reward arrives, it receives the basic factor α. Previously observed rewards move farther into the past, so their coefficients gain additional factors of 1 − α.

The Weighting Formula

If reward Rᵢ was observed n − i rewards ago, its weight is α(1 − α)ⁿ⁻ⁱ. The exponent n − i counts the number of rewards observed after Rᵢ.

weight of Rᵢ = α(1 − α)ⁿ⁻ⁱ

The initial estimate Q₁ also remains part of the weighted average. Its weight after n rewards is (1 − α)ⁿ. Thus, the current estimate contains a contribution from the initial estimate as well as contributions from the observed rewards.

contributes to current estimatecontributes to current estimateincreasing recencyQ₁(1 − α)ⁿR₁α(1 − α)ⁿ⁻¹Rᵢα(1 − α)ⁿ⁻ⁱRₙα
How does the current estimate contain the initial estimate and each past reward, and what coefficient multiplies each one?
Qₙ₊₁ = (1 − α)ⁿQ₁ + Σ from i = 1 to n of α(1 − α)ⁿ⁻ⁱRᵢ

All of these weights add to 1: (1 − α)ⁿ + Σ from i = 1 to n of α(1 − α)ⁿ⁻ⁱ = 1. That is why the result is a weighted average: the total contribution is distributed between the initial estimate and the observed rewards.

A Weight Calculation

Three rewards with α = 0.5

Suppose three rewards have been observed and α = 0.5. Find the coefficient of the initial estimate Q₁ and the coefficients of R₁, R₂, and R₃ in the current estimate.

Initial estimate: There are n = 3 observed rewards, so the coefficient of Q₁ is (1 − 0.5)³ = 0.5³ = 0.125.

First reward: R₁ was followed by two rewards, so its coefficient is 0.5(1 − 0.5)² = 0.5 × 0.25 = 0.125.

Second reward: R₂ was followed by one reward, so its coefficient is 0.5(1 − 0.5) = 0.25.

Third reward: R₃ is the most recent reward, so its coefficient is α = 0.5.

Check the total: The coefficients total 0.125 + 0.125 + 0.25 + 0.5 = 1, confirming that these contributions form a weighted average.

The coefficients are Q₁: 0.125, R₁: 0.125, R₂: 0.25, and R₃: 0.5. The most recent reward has the largest reward coefficient.

multiply by 1 − αmultiply by 1 − αmultiply by 1 − αRᵢαRᵢα(1 − α)Rᵢα(1 − α)²Rᵢα(1 − α)³
What happens to a particular reward's weight each time another reward is observed after it?

The Role of α

The constant step-size parameter α determines the basic weight assigned to a newly observed reward. In the weighting formula, the most recent reward has coefficient α. The previous information is reduced through factors of 1 − α. Therefore, α sets the balance between the new reward's direct contribution and the contribution retained from earlier information.

Part of the weightingCoefficientMeaning
Most recent rewardαThe new reward's direct weight
A reward one step olderα(1 − α)The reward's direct weight multiplied by one decay factor
Initial estimate after n rewards(1 − α)ⁿThe retained influence of the starting estimate

Why Recency Helps Adaptation

When the reward distribution changes, an estimate based equally on all past rewards may allow distant history to retain too much influence. Exponential recency weighting instead emphasizes rewards from the more recent situation while preserving progressively smaller contributions from earlier observations. This makes the weighting pattern suited to nonstationary problems.

The method adapts through repeated updates: every newly observed reward receives the direct factor α, while every older reward acquires another factor of 1 − α. The estimate therefore shifts toward recent evidence without eliminating the past immediately.

Common Weighting Mistakes

  • Giving every reward the same weight

    The method is specifically designed to give greater weight to recent rewards and smaller weight to older rewards.

    Fix: Count the number of intervening rewards and use α(1 − α)ⁿ⁻ⁱ.

  • Forgetting the initial estimate

    The initial estimate has coefficient (1 − α)ⁿ and remains part of the weighted average.

    Fix: Include the initial-estimate term before checking that all coefficients add to 1.

  • Using the number of elapsed time steps instead of intervening rewards

    The exponent n − i counts the intervening rewards.

    Fix: Start at the index of Rᵢ, identify the current reward index n, and calculate n − i.

  • Subtracting a fixed amount for each older reward

    The decrease is exponential: each additional intervening reward multiplies the previous weight by 1 − α.

    Fix: Apply another factor of 1 − α for every additional intervening reward.

Check Your Understanding

MEDIUM

Let α = 0.5 and suppose four rewards have been observed. What coefficient belongs to the first reward R₁? What coefficient belongs to the most recent reward R₄? What coefficient belongs to the initial estimate Q₁?

Hints
  • For R₁, count the three rewards observed after it.
  • The most recent reward has coefficient α.
  • The initial estimate's coefficient is (1 − α)ⁿ, with n = 4.

What do you think happens?

With α = 0.5 and four observed rewards, what are the three requested coefficients?

Reveal answer

Answer: R₁ has coefficient 0.5(0.5)³ = 0.0625. R₄ has coefficient 0.5. Q₁ has coefficient (0.5)⁴ = 0.0625.

R₁ has three intervening rewards, so its coefficient contains three factors of 1 − α. R₄ is the most recent reward, so its coefficient is α. The initial estimate receives one factor of 1 − α for each of the four rewards.

Key Takeaways

  1. A nonstationary problem can make old rewards less useful for estimating the present.
  2. The weight of reward Rᵢ is α(1 − α)ⁿ⁻ⁱ, where n − i is the number of intervening rewards.
  3. Each additional intervening reward multiplies an older reward's weight by 1 − α, producing exponential decay.
  4. The initial estimate Q₁ has weight (1 − α)ⁿ.
  5. All coefficients add to 1, so the estimate is a weighted average that emphasizes recent rewards.

Key Takeaways

  • Recent rewards receive greater influence than older rewards.
  • The coefficient of Rᵢ is α(1 − α)ⁿ⁻ⁱ, determined by the number of rewards observed after it.
  • The initial estimate contributes with coefficient (1 − α)ⁿ.
  • The factor 1 − α is applied once for every intervening reward, creating exponential decay.
  • The coefficients add to 1, making the result a weighted average.