Average Reward Methods
The average reward objective evaluates average reward over a policy's state distribution.
Start with the Policy
Average reward methods begin with a policy. That policy determines the distribution with which states occur under it. The average reward objective then evaluates the average reward over that policy-dependent state distribution.
Tracing the State Distribution
Comparing Two Policies through Their Distributions
Consider two policies whose states occur with different policy-dependent distributions. Explain what the average reward objective evaluates.
Identify the policy: Treat each policy as the starting point of the evaluation.
Identify its state distribution: Each policy determines the distribution with which states occur under that policy.
Evaluate the rewards over that distribution: The average reward objective evaluates average reward over the distribution associated with the policy being considered.
Compare the policies: The comparison concerns how the objective ranks the policies, rather than a calculation of particular numerical rewards.
The average reward objective evaluates each policy using the average reward over that policy's own state distribution.
The distribution is not an unrelated extra input. It is policy-dependent: start with a policy, determine the state distribution under that policy, and evaluate average reward over that distribution.
Adding Discounted Values
The proposed discounted objective follows the same policy-dependent state distribution, but it does not combine rewards in exactly the same way as the average reward description. It sums discounted values over the distribution with which states occur under the policy.
Following the Proposed Construction
Describe the evaluation sequence for one policy under the proposed discounted objective.
Begin with the policy: The policy determines the distribution with which states occur.
Keep that distribution: The proposed objective uses this same policy-dependent distribution rather than replacing it with an unrelated distribution.
Use state values: Associate values with the states represented in the distribution.
Apply discounting: Combine those values using discounting.
Form the objective: The proposed discounted objective sums the discounted values over the policy-dependent distribution.
The proposed objective combines two ingredients: the values associated with states and the distribution with which those states occur under the policy.
Comparing Policy Rankings
The central result is about ordering policies. Although the proposed discounted objective is written using discounted values summed over a policy-dependent state distribution, it orders policies identically to the undiscounted, or average reward, objective.
Why Discounting Is Not Rescued
The proposed construction changes how the objective is presented: it combines discounted values with the state distribution induced by the policy. However, this change does not produce a different policy ordering from the undiscounted average reward objective. Therefore, the proposal does not rescue discounting.
Using the policy's state distribution does not by itself make the proposed discounted objective rank policies differently. The source result is that its policy ordering remains identical to the average reward objective.
Common Reasoning Mistakes
Assuming that a new-looking objective must produce a new policy ranking.
The source result states that it orders policies identically to the undiscounted average reward objective.
Fix:
Evaluate an objective by the policy ordering it produces, not only by the way its expression is written.Treating the state distribution as independent of the policy.
The policy determines the distribution with which states occur under it.
Fix:
For each policy, first identify its policy-dependent state distribution.Describing the proposed objective as ordinary average reward without mentioning values or discounting.
The proposed objective sums discounted values over the policy-dependent distribution.
Fix:
Mention both ingredients: the state distribution under the policy and the discounted values associated with those states.Claiming that the proposal rescues discounting because it incorporates the state distribution.
The proposal does not produce a different policy ordering from the average reward objective.
Fix:
State the result precisely: the proposed discounted objective does not rescue discounting.
Check Your Understanding
A learner says: “The proposed discounted objective must rank policies differently because it uses discounted values.” Explain why this conclusion does not follow from the objective's construction.
Hints
- Identify the distribution used by the proposed objective.
- Ask what the source says about the ordering of policies.
- Distinguish an objective's written form from the policy ranking it produces.
What do you think happens?
If the proposed discounted objective and the average reward objective are compared by the policies they prefer, should their policy rankings agree or differ?
Reveal answer
Answer: They should agree.
The proposed discounted objective orders policies identically to the undiscounted average reward objective.
Key Takeaways
- The average reward objective evaluates average reward over the state distribution determined by a policy.
- The proposed discounted objective sums discounted values over that same policy-dependent state distribution.
- The important comparison is the policy ranking produced by each objective.
- The proposed discounted and undiscounted average reward objectives order policies identically.
- Because the policy ordering does not change, the proposed discounted objective does not rescue discounting.