Concepts / Average Reward Formulation

Average Reward Formulation

Differential value functions are the average-reward counterparts of familiar value functions.

  • Programming

Why Continuing Tasks Need a Different Objective

In an episodic problem, an agent can organize learning around episodes that eventually end. Continuing control is different: the interaction continues indefinitely, so the formulation must evaluate performance without depending on a terminal state. The average reward formulation addresses this setting by changing the objective to maximizing average reward per time step.

actscontinuescontinuesPolicy πcontinuing behaviorTime step 1rewardTime step 2rewardContinuing time stepsaverage reward per timestep
How does average reward evaluate an indefinitely continuing policy without requiring a terminal state?

Differential Values as Average-Reward Counterparts

The average-reward setting does not simply reuse every object from the usual formulation unchanged. It introduces differential versions of value functions, Bellman equations, and temporal-difference errors. These are parallel versions of familiar reinforcement-learning components: the conceptual changes are small, but the continuing average-reward setting requires its own corresponding definitions and algorithms.

A differential value function is an average-reward counterpart of a familiar value function. Its role belongs to the average-reward formulation, where the policy is evaluated in relation to average reward per time step rather than only through the usual formulation.

corresponds tocorresponds tocorresponds toFamiliar valuefunctionsusual formulationDifferential valuefunctionsaverage-reward counterpartBellman equationsusual formulationDifferential Bellmanequationsaverage-reward counterpartTD errorsusual formulationDifferential TDerrorsaverage-reward counterpart
How do differential value functions and related learning components correspond to their familiar counterparts?

The word differential signals a change of formulation, not the disappearance of familiar reinforcement-learning ideas. Value functions, Bellman equations, and temporal-difference errors still have parallel roles, but they are defined for the average-reward setting.

Semi-Gradient Sarsa with Approximate Values

Semi-gradient Sarsa becomes important when a value function is represented approximately rather than stored as a fully explicit table. Instead of treating every value as an independently listed quantity, the learner works with a parameterized approximation. The source identifies semi-gradient Sarsa with function approximation as a line of work first explored by Rummery and Niranjan in 1994. In the average-reward setting, this idea is developed through a parallel family of differential algorithms, including differential versions of semi-gradient n-step Sarsa.

evaluatessupportscontributesdriveschangesTransition experiencestate, action, andsubsequent experienceApproximate valueparameterized estimateSarsa targetlearning targetParameter updatesemi-gradient stepUpdated estimateapproximate value
How does semi-gradient Sarsa use an approximate value function to form a learning target and update its parameterized estimate?

The important connection is structural. Function approximation supplies a parameterized representation of value, while semi-gradient Sarsa supplies a learning procedure that uses experience to improve that representation. In the average-reward version, the corresponding quantities are differential rather than ordinary, so the procedure belongs to the parallel family of differential algorithms.

Reading Bounded-Region Behavior

A particularly important convergence statement concerns linear semi-gradient Sarsa with ε-greedy action selection. The reported behavior is that the parameter estimates enter a bounded region near the best solution rather than converging in the usual sense. This means the learning process gets and remains near a desirable region, but it should not be described as settling permanently on one exact parameter vector.

entercontainsParameter estimatesoutside the best regionBounded regionnear the best solutionBest solutionreference pointParameter estimatesremain near the region
What does it mean for parameter estimates to enter and remain near a bounded region instead of converging to one fixed point?
DescriptionWhat it means
Ordinary convergenceThe estimates approach one fixed solution in the usual sense.
Bounded-region behaviorThe estimates enter a bounded region near the best solution without necessarily approaching one fixed point.

Separating Representation from Policy Ranking

Approximate control introduces a limitation that matters for the discounted formulation: when value functions are approximated, most policies cannot be represented by a value function. If policies cannot all be represented in that way, value-function representation alone cannot provide a common basis for comparing arbitrary policies.

may not fitmay not fitleavesPolicy Aarbitrary policyValue-functionrepresentationapproximate caseMost policies notrepresentedcomparison problemPolicy Barbitrary policy
What limitation prevents the discounted formulation from carrying its usual control approach directly into the approximate case?

The scalar η(π) is the average reward of policy π. The average reward provides a quantity that can be used to rank arbitrary policies, including policies that cannot be represented by an approximate value function.

assignsassignscomparescomparesPolicy π1differential valuesη(π1)scalar average rewardPolicy π2differential valuesη(π2)scalar average rewardPolicy orderingcompare scalar rewards
How can η(π) rank policies even when their differential value functions are represented approximately?

Generated example: Suppose two continuing policies are being considered, Policy A and Policy B. Their approximate differential value functions need not provide a representable value function for every possible policy. The average-reward formulation instead associates each policy with its scalar η value. Comparing those scalar values supplies a ranking, so the task is no longer dependent on representing most policies as explicit value functions.

Keep two roles separate. Differential value functions support value estimation in the average-reward setting. The scalar η(π) supports comparison and ranking of policies. Approximate representation addresses the first role; policy ranking addresses the second.

Common Interpretation Errors

  • Treating average reward as an unchanged copy of the usual formulation.

    The average-reward setting introduces differential versions of these components.

    Fix: Look for the parallel differential value functions, differential Bellman equations, and differential TD errors.

  • Calling bounded-region behavior ordinary convergence.

    The reported behavior is entry into a bounded region near the best solution, not convergence in the usual sense.

    Fix: State that the estimates enter and remain near a bounded region.

  • Assuming approximate value representation can rank every policy.

    In the approximate case, most policies cannot be represented by a value function.

    Fix: Use η(π), the scalar average reward, as the common policy-ranking quantity.

  • Explaining continuing control only with an episodic objective.

    Continuing cases require a formulation based on maximizing average reward per time step.

    Fix: Identify average reward per time step as the continuing-control objective.

Check Your Understanding

MEDIUM

Explain, in your own words, why the average-reward formulation needs both differential value functions and the scalar η(π). Your answer should distinguish the role of estimating values from the role of ranking arbitrary policies.

Hints
  • Start with the fact that continuing control is organized around average reward per time step.
  • Mention the differential counterparts of familiar value-learning components.
  • Explain why approximate value functions cannot represent most policies.
  • Finish by describing η(π) as a common scalar for policy comparison.
EASY

A learner says: “Because the parameter estimates enter a bounded region near the best solution, the algorithm has converged normally.” Correct this statement using the precise convergence behavior reported for linear semi-gradient Sarsa with ε-greedy action selection.

Hints
  • Focus on the difference between a region and one fixed point.
  • Use the phrase bounded region near the best solution.

Key Takeaways

  1. Continuing control motivates an objective based on maximizing average reward per time step.
  2. Differential value functions, differential Bellman equations, and differential TD errors are the average-reward counterparts of familiar components.
  3. Semi-gradient Sarsa connects approximate, parameterized value functions with experience-based control learning.
  4. For linear semi-gradient Sarsa with ε-greedy action selection, the reported behavior is entry into a bounded region near the best solution, not ordinary convergence to one fixed point.
  5. Approximate value functions cannot represent most policies, so η(π), the scalar average reward of policy π, provides a way to rank arbitrary policies.

Key Takeaways

  • Average reward is designed for continuing control and evaluates a policy through average reward per time step.
  • Differential value functions and related differential learning components parallel the familiar reinforcement-learning formulation.
  • Semi-gradient Sarsa is relevant when values are represented by parameters rather than an explicit table.
  • Entering a bounded region near the best solution is not the same as converging to one fixed point.
  • The scalar η(π) separates policy ranking from approximate value-function representation.