Average-Reward Methods in Continuing Problems
Discounting does not automatically create a meaningful preference in a continuing problem with no distinguished beginning or end.
When No Moment Comes First
Discounting is often assumed to create a meaningful preference between rewards that arrive at different times. That assumption becomes questionable in a continuing problem with no distinguished beginning or end. There may be no first time step, no final time step, and no specially important state. If the available information is only an ongoing sequence of rewards and actions, choosing one point as the beginning can introduce a preference that the problem itself did not provide.
Shifting the Starting Point
Imagine an endless reward sequence with no natural first reward and no natural last reward. A discounted return is calculated from a chosen starting point, so rewards receive different discount positions relative to that point. But in a continuing evaluation, the starting point can move through the sequence. A reward that is immediate relative to one starting point becomes one step delayed relative to another, then two steps delayed, and so on. Because no time step is special, every reward follows the same pattern across the collection of returns.
What do you think happens?
If the starting point moves forward through a continuing reward sequence, does one particular reward remain in the same discounted position?
Reveal answer
Answer: No, the reward moves through the positions 1, γ, γ², and so on.
Across the different returns used in the continuing evaluation, one particular reward appears once without discounting, once with one factor of γ, once with two factors of γ, and so on. The absence of a privileged time step makes this pattern symmetric across rewards.
Average Reward as the Baseline
The average-reward setting measures performance by averaging rewards over a long interval. This measure fits a continuing problem because it does not require a specially selected first reward or a final reward. It asks how much reward the ongoing sequence produces on average rather than assigning special importance to rewards merely because they are close to an arbitrarily chosen beginning.
A long-run reward average
A continuing reward sequence alternates between 2 and 4. What average-reward value does this sequence suggest over a long interval?
Identify the repeating rewards: The sequence contains a reward of 2 and a reward of 4 in each repeated pair.
Add the rewards in one pair: The pair contributes 2 + 4 = 6.
Average the pair: Dividing 6 by the two rewards gives an average reward of 3.
The long-run average reward is 3.
The Symmetry Calculation
Take one particular reward at time t. As the evaluation starting point moves through the continuing sequence, that reward appears once without discounting, once with one factor of γ, once with two factors of γ, and so on. Its total contribution across these positions is the geometric sum 1 + γ + γ² + γ³ + ... = 1 / (1 - γ). Because no time step is special, the same pattern applies to every reward in the sequence.
Average discounted return for policy π = η(π) / (1 - γ)
Applying the common factor
Suppose a policy has average reward η(π) = 3. Using γ = 0.5, what average discounted return does the symmetry result give?
Write the relationship: The average discounted return is η(π) / (1 - γ).
Substitute the values: Substitute η(π) = 3 and γ = 0.5, giving 3 / (1 - 0.5).
Evaluate the factor: The denominator is 0.5, so the result is 3 / 0.5 = 6.
The average discounted return is 6. The value is the average reward 3 scaled by the common factor 1 / (1 - 0.5).
Policy Rankings Stay Fixed
The factor 1 / (1 - γ) is common to every policy in this continuing setting. Thus, each policy's average discounted return is its average reward multiplied by the same factor. Multiplying every policy score by the same common factor does not change which policy has the higher score. Therefore, changing γ does not change the ordering of policies under this formulation.
What do you think happens?
In this continuing setting, if policy A has a higher average reward than policy B, can changing γ reverse their ranking?
Reveal answer
Answer: No, the common factor scales both policies equally.
The average discounted return for each policy is its average reward multiplied by 1 / (1 - γ). Since the factor is shared, the ordering of policies remains unchanged.
Mistakes in Interpretation
Assuming discounting must create a meaningful preference in every reinforcement-learning problem.
A continuing problem may have no first time step, no final time step, and no specially important state.
Fix:
Check whether the problem supplies a distinguished beginning or end before interpreting discounting as a meaningful preference.Thinking that average discounted return creates a new policy ordering in this continuing setting.
Every policy's average discounted return is its average reward multiplied by the same factor 1 / (1 - γ).
Fix:
Compare the average rewards first; the common factor leaves their ordering unchanged.Forgetting why the geometric factor applies to every reward.
As the evaluation window moves through the continuing sequence, each reward occupies the discounted positions 1, γ, γ², and so on.
Fix:
Use the symmetry argument: no time step is special, so every reward follows the same discounted-position pattern.
Check Your Reasoning
A continuing problem has two policies. Policy A has average reward η(A) and Policy B has average reward η(B), with η(A) greater than η(B). Explain whether there is any value of γ that changes the ranking under the average discounted-return formulation described in this article.
Hints
- Write the average discounted return for each policy.
- Identify which part of the expression is shared.
- Ask whether multiplying both values by the same factor changes their ordering.
Explain in your own words why a reward at time t contributes with factors 1, γ, γ², and so on when the evaluation starting point moves through a continuing sequence. Then connect that pattern to the factor 1 / (1 - γ).
Hints
- Start from the absence of a privileged time step.
- Track one fixed reward as the starting point shifts.
- Recognize the resulting geometric sum.
What to Remember
- In a continuing problem with no distinguished beginning or end, discounting does not automatically create a meaningful preference.
- Average-reward methods assess performance by averaging rewards over a long interval.
- By symmetry, one reward occupies the discounted positions 1, γ, γ², and so on, whose total is 1 / (1 - γ).
- The average discounted return for policy π is η(π) / (1 - γ).
- Because the factor is common to all policies, changing γ does not change their ranking in this continuing formulation.
Key Takeaways
- A continuing problem may have no naturally important first or final time step.
- Average reward evaluates the ongoing sequence over a long interval without selecting a special beginning.
- The symmetry argument gives average discounted return η(π) / (1 - γ).
- The common scaling factor leaves policy rankings unchanged, so changing γ has no effect on the ordering of policies in this setting.