Concepts / Discounted Returns and Policy Evaluation

Discounted Returns and Policy Evaluation

Discounting does not automatically create a meaningful preference in a continuing problem with no distinguished beginning or end.

  • Programming

When No Time Step Is Special

Discounting is often treated as if it automatically creates a meaningful preference between earlier and later rewards. That assumption becomes questionable in an approximate continuing problem: there may be no distinguished beginning, no final time step, and no specially important state. If evaluation observes only an ongoing reward sequence, moving the evaluation window through that sequence does not give any particular time step a privileged role.

can identifyevaluatesFinite episodebeginning and endFirst rewardContinuing processno privileged time stepOngoing rewardsno natural first or lastreward
What changes when there is no distinguished beginning or end from which to measure the importance of rewards?

Average Reward as the Baseline

In the average-reward setting, performance is measured by averaging rewards over a long interval. The evaluation focuses on the long-run reward sequence rather than assigning special importance to a first or final time step. For a policy π, let η(π) denote its average reward. This gives a single measure of how well the policy performs across the continuing process.

continuescontinuesincluded inaverage rewardsRewardtime tRewardtime t + 1Rewardtime t + 2Long intervalmany ongoing rewardsη(π)average reward
How are rewards accumulated over an ongoing process summarized by a single long-run performance measure?

A long-run comparison

Suppose two policies are evaluated only through their ongoing reward sequences. Policy A has average reward η(A), while policy B has average reward η(B). How should their continuing performance be compared?

Choose the evaluation measure: Use the average reward over a long interval because the continuing process has no distinguished beginning or end.

Compare the averages: Policy A is preferred by this measure when η(A) is greater than η(B).

The average-reward setting ranks policies by their long-run average rewards.

The Symmetry Across Starting Times

The key argument is a symmetry argument. Imagine an endless reward sequence with no natural first reward and no natural last reward. Now average discounted returns over the different possible starting times used in the continuing evaluation. Consider one particular reward at time t. As the evaluation window moves, that reward appears once in the undiscounted position, once one step later in the discounted position, once two steps later, and so on. Its contributions therefore include one copy multiplied by 1, one copy multiplied by γ, one copy multiplied by γ², and every later discounted position.

appears atappears atappears atappears atReward rₜone particular rewardrₜfactor 1rₜfactor γrₜfactor γ²rₜfactor γⁿ and beyond
How can averaging discounted returns over every possible starting time make each reward appear equally often at every relative time position?

1 + γ + γ² + γ³ + ... = 1 / (1 - γ)

From Average Reward to Discounted Return

The symmetry argument applies to every reward in the ongoing sequence, not just to the one reward selected for inspection. Since no time step is special, every reward is counted through the same sequence of relative discount factors. The average discounted return for policy π is therefore η(π) / (1 - γ). Discounting has not created a new long-run preference in this setting; it has scaled the average reward by a common factor.

applies same pattern to each rewardsums tomultiplied by factorscalesη(π)average reward1 + γ + γ² + ...relative discount positions1 / (1 - γ)geometric-sum factorη(π) / (1 - γ)average discounted return
How does the average reward flow through the discounted sum to produce the factor 1 / (1 - γ)?

Applying the common factor

Suppose a policy has average reward η(π) equal to 3, and consider a discount factor γ equal to 0.5. Use the continuing-task relationship to determine its average discounted return.

Write the relationship: The average discounted return is η(π) / (1 - γ).

Substitute the values: Substitute η(π) = 3 and γ = 0.5, giving 3 / (1 - 0.5).

Evaluate the factor: The denominator is 0.5, so the result is 6.

The average discounted return is 6. This number is a scaled version of the average reward, not a different policy preference.

Why Policy Rankings Stay Fixed

The factor 1 / (1 - γ) is common to the policies being compared. If policy A has a larger average reward than policy B, multiplying both average rewards by that same factor preserves their ordering. Therefore, in this continuing setting, changing γ does not change the formulation's ranking of policies. It changes the numerical scale of the average discounted returns, but not which policy is ranked higher.

multiply bymultiply bysame factorsame factorη(A)higher average rewardη(A) / (1 - γ)higher discounted returnη(B)lower average rewardη(B) / (1 - γ)lower discounted return1 / (1 - γ)common factor
If every policy's average discounted return is its average reward multiplied by the same factor, can changing γ change which policy ranks higher?

What do you think happens?

Policy A has a higher average reward than policy B. In this continuing setting, what happens to their ranking when γ changes?

  • Policy A and policy B may swap positions.
  • Their ranking stays the same, although their numerical returns are rescaled.
  • Both policies receive exactly the same return.
Reveal answer

Answer: Their ranking stays the same, although their numerical returns are rescaled.

Both policies' average rewards are multiplied by the same factor, 1 / (1 - γ). A common factor does not change their ordering.

Mistakes in Continuing Evaluation

  • Assuming that discounting must create a new preference in every reinforcement-learning problem.

    A continuing problem may have no first time step, no final time step, and no specially important state. Averaging discounted returns can then produce only a scaled version of the average reward.

    Fix: First ask whether the task has a distinguished beginning or end. For the continuing setting described here, use the symmetry argument and compare long-run average rewards.

  • Thinking that changing γ must change which policy is best.

    The average discounted return for each policy is its average reward multiplied by the same factor 1 / (1 - γ).

    Fix: Separate numerical scale from policy ordering. The values change with the common factor, but the ranking remains unchanged.

  • Tracking only one reward position instead of the full continuing sequence.

    The symmetry argument depends on averaging over possible starting times. Each reward appears with factors 1, γ, γ², γ³, and so on.

    Fix: Trace one reward across every relative position, then apply the same reasoning to all rewards.

Check the Symmetry

MEDIUM

A continuing problem has two policies. Policy A has average reward η(A) = 4, and policy B has average reward η(B) = 5. Without calculating exact discounted-return values, predict which policy is ranked higher under the continuing evaluation when γ changes. Then explain why the common factor does not reverse the ordering.

Hints
  • Write the average discounted return for each policy using η(π) / (1 - γ).
  • Identify whether the denominator is different for the two policies.
  • Compare the numerators after recognizing that both policies use the same γ.

Practice solution

Determine the ranking of the two policies from the practice question.

Write both returns: Policy A has average discounted return 4 / (1 - γ), while policy B has average discounted return 5 / (1 - γ).

Compare the shared factor: Both expressions contain the same factor 1 / (1 - γ), so the factor scales both values equally.

Preserve the ordering: Because 5 is greater than 4, policy B remains higher under the continuing evaluation.

Policy B ranks higher for every change in γ covered by this continuing formulation. Changing γ changes the numerical scale but not the policy ordering.

Key Takeaways

  1. A continuing problem may have no first time step, final time step, or specially important state, so discounting does not automatically create a meaningful preference.
  2. Average-reward evaluation measures performance by averaging rewards over a long interval.
  3. Across all possible starting times, each reward appears at every relative discounted position, producing the geometric factor 1 / (1 - γ).
  4. For policy π, the average discounted return is η(π) / (1 - γ).
  5. Because the factor is common across policies, changing γ changes the scale of returns but not their ranking in this continuing setting.

Key Takeaways

  • Discounting is questionable when an approximate continuing problem has no distinguished beginning, end, or time step.
  • The average-reward setting evaluates a policy by its long-run average reward η(π).
  • Symmetry across starting times makes every reward contribute through the geometric sum 1 + γ + γ² + ... = 1 / (1 - γ).
  • The average discounted return is η(π) / (1 - γ).
  • The common factor preserves policy rankings, so changing γ does not change the ordering in this continuing formulation.