Concepts / Average Reward Methods

Average Reward Methods

The average reward objective evaluates average reward over a policy's state distribution.

  • Programming

Start with the Policy

Average reward methods begin with a policy. That policy determines the distribution with which states occur under it. The average reward objective then evaluates the average reward over that policy-dependent state distribution.

determinesweights occurrence ofcombines intoPolicyState distributionunder the policyState rewardsrewards in occurring statesAverage rewardobjective value
How do rewards earned in different states combine according to the distribution of states visited under a policy?

Tracing the State Distribution

Comparing Two Policies through Their Distributions

Consider two policies whose states occur with different policy-dependent distributions. Explain what the average reward objective evaluates.

Identify the policy: Treat each policy as the starting point of the evaluation.

Identify its state distribution: Each policy determines the distribution with which states occur under that policy.

Evaluate the rewards over that distribution: The average reward objective evaluates average reward over the distribution associated with the policy being considered.

Compare the policies: The comparison concerns how the objective ranks the policies, rather than a calculation of particular numerical rewards.

The average reward objective evaluates each policy using the average reward over that policy's own state distribution.

The distribution is not an unrelated extra input. It is policy-dependent: start with a policy, determine the state distribution under that policy, and evaluate average reward over that distribution.

Adding Discounted Values

The proposed discounted objective follows the same policy-dependent state distribution, but it does not combine rewards in exactly the same way as the average reward description. It sums discounted values over the distribution with which states occur under the policy.

determinesprovides state weightingare combined usingcontributesPolicyState distributionunder the policyState valuesassociated with statesDiscountingapplied to valuesProposed objectivesum over the distribution
How do a policy's state values and its state distribution combine to produce the proposed discounted objective?

Following the Proposed Construction

Describe the evaluation sequence for one policy under the proposed discounted objective.

Begin with the policy: The policy determines the distribution with which states occur.

Keep that distribution: The proposed objective uses this same policy-dependent distribution rather than replacing it with an unrelated distribution.

Use state values: Associate values with the states represented in the distribution.

Apply discounting: Combine those values using discounting.

Form the objective: The proposed discounted objective sums the discounted values over the policy-dependent distribution.

The proposed objective combines two ingredients: the values associated with states and the distribution with which those states occur under the policy.

Comparing Policy Rankings

The central result is about ordering policies. Although the proposed discounted objective is written using discounted values summed over a policy-dependent state distribution, it orders policies identically to the undiscounted, or average reward, objective.

ranks aboveranks abovePolicy Aaverage reward: higherPolicy Aproposed objective: higherPolicy Baverage reward: lowerPolicy Bproposed objective: lower
When do the average reward and proposed discounted objectives rank policies in the same order?

Why Discounting Is Not Rescued

The proposed construction changes how the objective is presented: it combines discounted values with the state distribution induced by the policy. However, this change does not produce a different policy ordering from the undiscounted average reward objective. Therefore, the proposal does not rescue discounting.

producesproducesAverage rewardobjectiveundiscountedProposed discountedobjectivediscounted values over thedistributionPolicy orderingaverage reward orderingPolicy orderingsame ordering
Which effects of discounting remain after values are combined with the policy's state distribution?

Using the policy's state distribution does not by itself make the proposed discounted objective rank policies differently. The source result is that its policy ordering remains identical to the average reward objective.

Common Reasoning Mistakes

  • Assuming that a new-looking objective must produce a new policy ranking.

    The source result states that it orders policies identically to the undiscounted average reward objective.

    Fix: Evaluate an objective by the policy ordering it produces, not only by the way its expression is written.

  • Treating the state distribution as independent of the policy.

    The policy determines the distribution with which states occur under it.

    Fix: For each policy, first identify its policy-dependent state distribution.

  • Describing the proposed objective as ordinary average reward without mentioning values or discounting.

    The proposed objective sums discounted values over the policy-dependent distribution.

    Fix: Mention both ingredients: the state distribution under the policy and the discounted values associated with those states.

  • Claiming that the proposal rescues discounting because it incorporates the state distribution.

    The proposal does not produce a different policy ordering from the average reward objective.

    Fix: State the result precisely: the proposed discounted objective does not rescue discounting.

Check Your Understanding

MEDIUM

A learner says: “The proposed discounted objective must rank policies differently because it uses discounted values.” Explain why this conclusion does not follow from the objective's construction.

Hints
  • Identify the distribution used by the proposed objective.
  • Ask what the source says about the ordering of policies.
  • Distinguish an objective's written form from the policy ranking it produces.

What do you think happens?

If the proposed discounted objective and the average reward objective are compared by the policies they prefer, should their policy rankings agree or differ?

  • They should agree.
  • They should differ.
Reveal answer

Answer: They should agree.

The proposed discounted objective orders policies identically to the undiscounted average reward objective.

Key Takeaways

  • The average reward objective evaluates average reward over the state distribution determined by a policy.
  • The proposed discounted objective sums discounted values over that same policy-dependent state distribution.
  • The important comparison is the policy ranking produced by each objective.
  • The proposed discounted and undiscounted average reward objectives order policies identically.
  • Because the policy ordering does not change, the proposed discounted objective does not rescue discounting.