Limitations of the Discounted Setting
The average reward objective evaluates average reward over a policy's state distribution.
Why Policy Rankings Matter
Two objectives can be written in different ways yet still prefer policies in exactly the same order. That is the key issue here. The proposed discounted objective changes how evaluation is described, but it does not change which policies look better or worse than others when compared with the undiscounted average reward objective.
To evaluate whether the proposal solves a limitation, compare its policy rankings with the rankings produced by the average reward objective.
From Policy to Average Reward
Begin with a policy. The policy determines the distribution with which states occur under that policy. The average reward objective evaluates the average reward over this policy-dependent state distribution. In other words, the objective is tied to the states that the policy causes to occur and the rewards associated with those states.
Adding Discounted Values
The proposed discounted objective follows the same policy-induced state distribution. Its distinctive step is to combine the values associated with those states using discounting. Thus, the proposal does not replace the policy's state distribution with a separate one; it sums discounted values over the distribution with which states occur under the policy.
The important comparison is therefore not simply average reward versus a generic discounted quantity. It is average reward over a policy's state distribution versus discounted values summed over that same policy-dependent distribution.
Comparing the Two Evaluation Lenses
| Objective | Distribution used | Quantity combined over that distribution |
|---|---|---|
| Average reward objective | The distribution induced by the policy | Average reward |
| Proposed discounted objective | The distribution induced by the policy | Discounted values |
The Ranking Result
Comparing Two Policies
Suppose two policies, Policy A and Policy B, are evaluated using both the average reward objective and the proposed discounted objective. What should you compare?
Step 1: Follow each policy: Each policy determines the distribution with which states occur under that policy.
Step 2: Apply the average reward objective: Evaluate the average reward over each policy's own state distribution.
Step 3: Apply the proposed discounted objective: For each policy, sum discounted values over the state distribution induced by that same policy.
Step 4: Compare the rankings: The source result is that the proposed discounted objective and the undiscounted average reward objective agree in how they order policies.
The two objectives produce the same policy ordering. The proposed discounted formulation does not create a different preference ordering between Policy A and Policy B.
Different objective descriptions do not necessarily imply different decisions. Here, the proposed discounted objective orders policies identically to the undiscounted average reward objective.
Why Discounting Is Not Rescued
The proposed objective does not rescue discounting because its construction still combines state values using discounting. It changes the distribution over which those discounted values are summed, making that distribution the one induced by the policy. However, this change does not produce a different policy ordering from the average reward objective. The proposal therefore retains the relevant dependence on discounting without solving the limitation under discussion.
Common Interpretation Errors
Assuming that a discounted objective must rank policies differently from an average reward objective.
The source result is that it orders policies identically to the undiscounted average reward objective.
Fix:
Compare the policy rankings, not only the vocabulary or form used to describe each objective.Ignoring the policy-induced state distribution.
The distribution is determined by the policy, and the proposal sums discounted values over that policy-dependent distribution.
Fix:
For every policy, first identify that policy's state distribution, then describe how the objective evaluates quantities over it.Claiming that using the policy's state distribution removes discounting.
The proposed objective still combines the associated state values using discounting.
Fix:
State clearly that the proposal changes the distribution used in the formulation but does not eliminate discounting or alter the resulting policy ordering.
Check Your Understanding
Explain, in your own words, why the proposed discounted objective can use a policy-dependent state distribution and still fail to rescue discounting.
Hints
- State what determines the distribution of states.
- Identify what the average reward objective evaluates over that distribution.
- Identify what the proposed discounted objective combines over the same distribution.
- Finish by comparing the policy rankings produced by the two objectives.
What do you think happens?
If the proposed discounted objective and the average reward objective are applied to the same set of policies, should you expect them to produce different policy orderings?
Reveal answer
Answer: No, the proposed discounted objective and the undiscounted average reward objective order policies identically.
The proposed objective sums discounted values over the state distribution induced by each policy, but this formulation does not produce a different policy ordering.
Core Takeaways
- A policy determines the distribution with which states occur under that policy.
- The average reward objective evaluates average reward over the policy's state distribution.
- The proposed discounted objective sums discounted values over that same policy-dependent distribution.
- The two objectives order policies identically.
- Because the proposed objective retains discounting and does not change policy rankings, it does not rescue discounting.
Key Takeaways
- The average reward objective evaluates average reward over a policy's state distribution.
- The proposed discounted objective combines discounted state values over the distribution induced by the policy.
- The proposed objective and the undiscounted average reward objective produce the same policy ordering.
- Using a policy-dependent distribution does not remove the proposed objective's dependence on discounting.
- The proposal therefore does not rescue discounting.