Average Reward Formulation
Differential value functions are the average-reward counterparts of familiar value functions.
Why Continuing Tasks Need a Different Objective
In an episodic problem, an agent can organize learning around episodes that eventually end. Continuing control is different: the interaction continues indefinitely, so the formulation must evaluate performance without depending on a terminal state. The average reward formulation addresses this setting by changing the objective to maximizing average reward per time step.
Differential Values as Average-Reward Counterparts
The average-reward setting does not simply reuse every object from the usual formulation unchanged. It introduces differential versions of value functions, Bellman equations, and temporal-difference errors. These are parallel versions of familiar reinforcement-learning components: the conceptual changes are small, but the continuing average-reward setting requires its own corresponding definitions and algorithms.
A differential value function is an average-reward counterpart of a familiar value function. Its role belongs to the average-reward formulation, where the policy is evaluated in relation to average reward per time step rather than only through the usual formulation.
The word differential signals a change of formulation, not the disappearance of familiar reinforcement-learning ideas. Value functions, Bellman equations, and temporal-difference errors still have parallel roles, but they are defined for the average-reward setting.
Semi-Gradient Sarsa with Approximate Values
Semi-gradient Sarsa becomes important when a value function is represented approximately rather than stored as a fully explicit table. Instead of treating every value as an independently listed quantity, the learner works with a parameterized approximation. The source identifies semi-gradient Sarsa with function approximation as a line of work first explored by Rummery and Niranjan in 1994. In the average-reward setting, this idea is developed through a parallel family of differential algorithms, including differential versions of semi-gradient n-step Sarsa.
The important connection is structural. Function approximation supplies a parameterized representation of value, while semi-gradient Sarsa supplies a learning procedure that uses experience to improve that representation. In the average-reward version, the corresponding quantities are differential rather than ordinary, so the procedure belongs to the parallel family of differential algorithms.
Reading Bounded-Region Behavior
A particularly important convergence statement concerns linear semi-gradient Sarsa with ε-greedy action selection. The reported behavior is that the parameter estimates enter a bounded region near the best solution rather than converging in the usual sense. This means the learning process gets and remains near a desirable region, but it should not be described as settling permanently on one exact parameter vector.
| Description | What it means |
|---|---|
| Ordinary convergence | The estimates approach one fixed solution in the usual sense. |
| Bounded-region behavior | The estimates enter a bounded region near the best solution without necessarily approaching one fixed point. |
Separating Representation from Policy Ranking
Approximate control introduces a limitation that matters for the discounted formulation: when value functions are approximated, most policies cannot be represented by a value function. If policies cannot all be represented in that way, value-function representation alone cannot provide a common basis for comparing arbitrary policies.
The scalar η(π) is the average reward of policy π. The average reward provides a quantity that can be used to rank arbitrary policies, including policies that cannot be represented by an approximate value function.
Generated example: Suppose two continuing policies are being considered, Policy A and Policy B. Their approximate differential value functions need not provide a representable value function for every possible policy. The average-reward formulation instead associates each policy with its scalar η value. Comparing those scalar values supplies a ranking, so the task is no longer dependent on representing most policies as explicit value functions.
Keep two roles separate. Differential value functions support value estimation in the average-reward setting. The scalar η(π) supports comparison and ranking of policies. Approximate representation addresses the first role; policy ranking addresses the second.
Common Interpretation Errors
Treating average reward as an unchanged copy of the usual formulation.
The average-reward setting introduces differential versions of these components.
Fix:
Look for the parallel differential value functions, differential Bellman equations, and differential TD errors.Calling bounded-region behavior ordinary convergence.
The reported behavior is entry into a bounded region near the best solution, not convergence in the usual sense.
Fix:
State that the estimates enter and remain near a bounded region.Assuming approximate value representation can rank every policy.
In the approximate case, most policies cannot be represented by a value function.
Fix:
Use η(π), the scalar average reward, as the common policy-ranking quantity.Explaining continuing control only with an episodic objective.
Continuing cases require a formulation based on maximizing average reward per time step.
Fix:
Identify average reward per time step as the continuing-control objective.
Check Your Understanding
Explain, in your own words, why the average-reward formulation needs both differential value functions and the scalar η(π). Your answer should distinguish the role of estimating values from the role of ranking arbitrary policies.
Hints
- Start with the fact that continuing control is organized around average reward per time step.
- Mention the differential counterparts of familiar value-learning components.
- Explain why approximate value functions cannot represent most policies.
- Finish by describing η(π) as a common scalar for policy comparison.
A learner says: “Because the parameter estimates enter a bounded region near the best solution, the algorithm has converged normally.” Correct this statement using the precise convergence behavior reported for linear semi-gradient Sarsa with ε-greedy action selection.
Hints
- Focus on the difference between a region and one fixed point.
- Use the phrase bounded region near the best solution.
Key Takeaways
- Continuing control motivates an objective based on maximizing average reward per time step.
- Differential value functions, differential Bellman equations, and differential TD errors are the average-reward counterparts of familiar components.
- Semi-gradient Sarsa connects approximate, parameterized value functions with experience-based control learning.
- For linear semi-gradient Sarsa with ε-greedy action selection, the reported behavior is entry into a bounded region near the best solution, not ordinary convergence to one fixed point.
- Approximate value functions cannot represent most policies, so η(π), the scalar average reward of policy π, provides a way to rank arbitrary policies.
Key Takeaways
- Average reward is designed for continuing control and evaluates a policy through average reward per time step.
- Differential value functions and related differential learning components parallel the familiar reinforcement-learning formulation.
- Semi-gradient Sarsa is relevant when values are represented by parameters rather than an explicit table.
- Entering a bounded region near the best solution is not the same as converging to one fixed point.
- The scalar η(π) separates policy ranking from approximate value-function representation.