Discounted Returns and Policy Evaluation
Discounting does not automatically create a meaningful preference in a continuing problem with no distinguished beginning or end.
When No Time Step Is Special
Discounting is often treated as if it automatically creates a meaningful preference between earlier and later rewards. That assumption becomes questionable in an approximate continuing problem: there may be no distinguished beginning, no final time step, and no specially important state. If evaluation observes only an ongoing reward sequence, moving the evaluation window through that sequence does not give any particular time step a privileged role.
Average Reward as the Baseline
In the average-reward setting, performance is measured by averaging rewards over a long interval. The evaluation focuses on the long-run reward sequence rather than assigning special importance to a first or final time step. For a policy π, let η(π) denote its average reward. This gives a single measure of how well the policy performs across the continuing process.
A long-run comparison
Suppose two policies are evaluated only through their ongoing reward sequences. Policy A has average reward η(A), while policy B has average reward η(B). How should their continuing performance be compared?
Choose the evaluation measure: Use the average reward over a long interval because the continuing process has no distinguished beginning or end.
Compare the averages: Policy A is preferred by this measure when η(A) is greater than η(B).
The average-reward setting ranks policies by their long-run average rewards.
The Symmetry Across Starting Times
The key argument is a symmetry argument. Imagine an endless reward sequence with no natural first reward and no natural last reward. Now average discounted returns over the different possible starting times used in the continuing evaluation. Consider one particular reward at time t. As the evaluation window moves, that reward appears once in the undiscounted position, once one step later in the discounted position, once two steps later, and so on. Its contributions therefore include one copy multiplied by 1, one copy multiplied by γ, one copy multiplied by γ², and every later discounted position.
1 + γ + γ² + γ³ + ... = 1 / (1 - γ)
From Average Reward to Discounted Return
The symmetry argument applies to every reward in the ongoing sequence, not just to the one reward selected for inspection. Since no time step is special, every reward is counted through the same sequence of relative discount factors. The average discounted return for policy π is therefore η(π) / (1 - γ). Discounting has not created a new long-run preference in this setting; it has scaled the average reward by a common factor.
Applying the common factor
Suppose a policy has average reward η(π) equal to 3, and consider a discount factor γ equal to 0.5. Use the continuing-task relationship to determine its average discounted return.
Write the relationship: The average discounted return is η(π) / (1 - γ).
Substitute the values: Substitute η(π) = 3 and γ = 0.5, giving 3 / (1 - 0.5).
Evaluate the factor: The denominator is 0.5, so the result is 6.
The average discounted return is 6. This number is a scaled version of the average reward, not a different policy preference.
Why Policy Rankings Stay Fixed
The factor 1 / (1 - γ) is common to the policies being compared. If policy A has a larger average reward than policy B, multiplying both average rewards by that same factor preserves their ordering. Therefore, in this continuing setting, changing γ does not change the formulation's ranking of policies. It changes the numerical scale of the average discounted returns, but not which policy is ranked higher.
What do you think happens?
Policy A has a higher average reward than policy B. In this continuing setting, what happens to their ranking when γ changes?
Reveal answer
Answer: Their ranking stays the same, although their numerical returns are rescaled.
Both policies' average rewards are multiplied by the same factor, 1 / (1 - γ). A common factor does not change their ordering.
Mistakes in Continuing Evaluation
Assuming that discounting must create a new preference in every reinforcement-learning problem.
A continuing problem may have no first time step, no final time step, and no specially important state. Averaging discounted returns can then produce only a scaled version of the average reward.
Fix:
First ask whether the task has a distinguished beginning or end. For the continuing setting described here, use the symmetry argument and compare long-run average rewards.Thinking that changing γ must change which policy is best.
The average discounted return for each policy is its average reward multiplied by the same factor 1 / (1 - γ).
Fix:
Separate numerical scale from policy ordering. The values change with the common factor, but the ranking remains unchanged.Tracking only one reward position instead of the full continuing sequence.
The symmetry argument depends on averaging over possible starting times. Each reward appears with factors 1, γ, γ², γ³, and so on.
Fix:
Trace one reward across every relative position, then apply the same reasoning to all rewards.
Check the Symmetry
A continuing problem has two policies. Policy A has average reward η(A) = 4, and policy B has average reward η(B) = 5. Without calculating exact discounted-return values, predict which policy is ranked higher under the continuing evaluation when γ changes. Then explain why the common factor does not reverse the ordering.
Hints
- Write the average discounted return for each policy using η(π) / (1 - γ).
- Identify whether the denominator is different for the two policies.
- Compare the numerators after recognizing that both policies use the same γ.
Practice solution
Determine the ranking of the two policies from the practice question.
Write both returns: Policy A has average discounted return 4 / (1 - γ), while policy B has average discounted return 5 / (1 - γ).
Compare the shared factor: Both expressions contain the same factor 1 / (1 - γ), so the factor scales both values equally.
Preserve the ordering: Because 5 is greater than 4, policy B remains higher under the continuing evaluation.
Policy B ranks higher for every change in γ covered by this continuing formulation. Changing γ changes the numerical scale but not the policy ordering.
Key Takeaways
- A continuing problem may have no first time step, final time step, or specially important state, so discounting does not automatically create a meaningful preference.
- Average-reward evaluation measures performance by averaging rewards over a long interval.
- Across all possible starting times, each reward appears at every relative discounted position, producing the geometric factor 1 / (1 - γ).
- For policy π, the average discounted return is η(π) / (1 - γ).
- Because the factor is common across policies, changing γ changes the scale of returns but not their ranking in this continuing setting.
Key Takeaways
- Discounting is questionable when an approximate continuing problem has no distinguished beginning, end, or time step.
- The average-reward setting evaluates a policy by its long-run average reward η(π).
- Symmetry across starting times makes every reward contribute through the geometric sum 1 + γ + γ² + ... = 1 / (1 - γ).
- The average discounted return is η(π) / (1 - γ).
- The common factor preserves policy rankings, so changing γ does not change the ordering in this continuing formulation.