Concepts / Average-Reward Methods in Continuing Problems

Average-Reward Methods in Continuing Problems

Discounting does not automatically create a meaningful preference in a continuing problem with no distinguished beginning or end.

  • Programming

When No Moment Comes First

Discounting is often assumed to create a meaningful preference between rewards that arrive at different times. That assumption becomes questionable in a continuing problem with no distinguished beginning or end. There may be no first time step, no final time step, and no specially important state. If the available information is only an ongoing sequence of rewards and actions, choosing one point as the beginning can introduce a preference that the problem itself did not provide.

Shifting the Starting Point

Imagine an endless reward sequence with no natural first reward and no natural last reward. A discounted return is calculated from a chosen starting point, so rewards receive different discount positions relative to that point. But in a continuing evaluation, the starting point can move through the sequence. A reward that is immediate relative to one starting point becomes one step delayed relative to another, then two steps delayed, and so on. Because no time step is special, every reward follows the same pattern across the collection of returns.

next rewardnext rewardnext rewardnext rewardr₀immediater₀γr₁γr₁γ²r₂γ²r₂γ³
What changes when the same continuing reward stream is viewed from different starting points, and why does discounting introduce an arbitrary preference?

What do you think happens?

If the starting point moves forward through a continuing reward sequence, does one particular reward remain in the same discounted position?

  • Yes, every reward always has the same discount factor
  • No, the reward moves through the positions 1, γ, γ², and so on
  • Only the first reward changes position
Reveal answer

Answer: No, the reward moves through the positions 1, γ, γ², and so on.

Across the different returns used in the continuing evaluation, one particular reward appears once without discounting, once with one factor of γ, once with two factors of γ, and so on. The absence of a privileged time step makes this pattern symmetric across rewards.

Average Reward as the Baseline

The average-reward setting measures performance by averaging rewards over a long interval. This measure fits a continuing problem because it does not require a specially selected first reward or a final reward. It asks how much reward the ongoing sequence produces on average rather than assigning special importance to rewards merely because they are close to an arbitrarily chosen beginning.

A long-run reward average

A continuing reward sequence alternates between 2 and 4. What average-reward value does this sequence suggest over a long interval?

Identify the repeating rewards: The sequence contains a reward of 2 and a reward of 4 in each repeated pair.

Add the rewards in one pair: The pair contributes 2 + 4 = 6.

Average the pair: Dividing 6 by the two rewards gives an average reward of 3.

The long-run average reward is 3.

continuescontinuescontinueslong-run averagelong-run averagelong-run averagelong-run average2reward3average reward4reward2reward4reward
How do rewards accumulate over a continuing sequence, and how is long-run average reward computed without relying on a special beginning or end?

The Symmetry Calculation

Take one particular reward at time t. As the evaluation starting point moves through the continuing sequence, that reward appears once without discounting, once with one factor of γ, once with two factors of γ, and so on. Its total contribution across these positions is the geometric sum 1 + γ + γ² + γ³ + ... = 1 / (1 - γ). Because no time step is special, the same pattern applies to every reward in the sequence.

Average discounted return for policy π = η(π) / (1 - γ)

Applying the common factor

Suppose a policy has average reward η(π) = 3. Using γ = 0.5, what average discounted return does the symmetry result give?

Write the relationship: The average discounted return is η(π) / (1 - γ).

Substitute the values: Substitute η(π) = 3 and γ = 0.5, giving 3 / (1 - 0.5).

Evaluate the factor: The denominator is 0.5, so the result is 3 / 0.5 = 6.

The average discounted return is 6. The value is the average reward 3 scaled by the common factor 1 / (1 - 0.5).

next starting pointnext starting pointnext starting pointsum of all positionssum of all positionssum of all positionssum of all positionsrₜ11 / (1 - γ)total contributionrₜγrₜγ²rₜγ³
How does shifting the starting point through a continuing reward sequence show that the average discounted return equals the average reward multiplied by 1 / (1 - γ)?

Policy Rankings Stay Fixed

The factor 1 / (1 - γ) is common to every policy in this continuing setting. Thus, each policy's average discounted return is its average reward multiplied by the same factor. Multiplying every policy score by the same common factor does not change which policy has the higher score. Therefore, changing γ does not change the ordering of policies under this formulation.

multiplymultiplysame scalingsame scalingPolicy Aaverage reward η(A)1 / (1 - γ)common factorPolicy Adiscounted rankPolicy Baverage reward η(B)Policy Bdiscounted rank
If every policy's average discounted return is its average reward scaled by the same factor 1 / (1 - γ), does changing γ alter which policy ranks higher?

What do you think happens?

In this continuing setting, if policy A has a higher average reward than policy B, can changing γ reverse their ranking?

  • Yes, changing γ can reverse the ranking
  • No, the common factor scales both policies equally
  • Only if the continuing sequence has no actions
Reveal answer

Answer: No, the common factor scales both policies equally.

The average discounted return for each policy is its average reward multiplied by 1 / (1 - γ). Since the factor is shared, the ordering of policies remains unchanged.

Mistakes in Interpretation

  • Assuming discounting must create a meaningful preference in every reinforcement-learning problem.

    A continuing problem may have no first time step, no final time step, and no specially important state.

    Fix: Check whether the problem supplies a distinguished beginning or end before interpreting discounting as a meaningful preference.

  • Thinking that average discounted return creates a new policy ordering in this continuing setting.

    Every policy's average discounted return is its average reward multiplied by the same factor 1 / (1 - γ).

    Fix: Compare the average rewards first; the common factor leaves their ordering unchanged.

  • Forgetting why the geometric factor applies to every reward.

    As the evaluation window moves through the continuing sequence, each reward occupies the discounted positions 1, γ, γ², and so on.

    Fix: Use the symmetry argument: no time step is special, so every reward follows the same discounted-position pattern.

Check Your Reasoning

MEDIUM

A continuing problem has two policies. Policy A has average reward η(A) and Policy B has average reward η(B), with η(A) greater than η(B). Explain whether there is any value of γ that changes the ranking under the average discounted-return formulation described in this article.

Hints
  • Write the average discounted return for each policy.
  • Identify which part of the expression is shared.
  • Ask whether multiplying both values by the same factor changes their ordering.
MEDIUM

Explain in your own words why a reward at time t contributes with factors 1, γ, γ², and so on when the evaluation starting point moves through a continuing sequence. Then connect that pattern to the factor 1 / (1 - γ).

Hints
  • Start from the absence of a privileged time step.
  • Track one fixed reward as the starting point shifts.
  • Recognize the resulting geometric sum.

What to Remember

  1. In a continuing problem with no distinguished beginning or end, discounting does not automatically create a meaningful preference.
  2. Average-reward methods assess performance by averaging rewards over a long interval.
  3. By symmetry, one reward occupies the discounted positions 1, γ, γ², and so on, whose total is 1 / (1 - γ).
  4. The average discounted return for policy π is η(π) / (1 - γ).
  5. Because the factor is common to all policies, changing γ does not change their ranking in this continuing formulation.

Key Takeaways

  • A continuing problem may have no naturally important first or final time step.
  • Average reward evaluates the ongoing sequence over a long interval without selecting a special beginning.
  • The symmetry argument gives average discounted return η(π) / (1 - γ).
  • The common scaling factor leaves policy rankings unchanged, so changing γ has no effect on the ordering of policies in this setting.