Sarsa and n-step Sarsa
Differential value functions are the average-reward counterparts of familiar value functions.
Two Ways to Read Learning Progress
When reinforcement learning is studied with an average-reward objective, the familiar value-function vocabulary needs a parallel version. The source describes differential value functions as the average-reward counterparts of familiar value functions. This is not a cosmetic change: the average-reward setting introduces corresponding differential value functions, differential Bellman equations, and differential temporal-difference errors.
The same parallel idea applies to Sarsa. Semi-gradient Sarsa is useful when the action-value function is represented approximately rather than stored as a fully explicit table. With linear semi-gradient Sarsa and ε-greedy action selection, the reported behavior is especially important: the parameters enter a bounded region near the best solution instead of converging in the usual sense to one fixed point.
The Differential Perspective
Average-reward reinforcement learning does not simply reuse every object from the usual formulation unchanged. Instead, it develops corresponding differential objects. Differential value functions are the value-function counterpart for the average-reward setting, and differential Bellman equations and temporal-difference errors form a parallel set of components. The organizing principle is continuity: the conceptual changes are small enough to form a parallel family, but the average-reward setting still requires its own definitions and algorithms.
A useful reading strategy is to ask which familiar component is being adapted: value functions have differential counterparts, Bellman equations have differential counterparts, and temporal-difference errors have differential counterparts.
Generated example: Imagine comparing two continuing learning situations by their long-run reward behavior rather than by treating each situation as a separate finite episode. The differential vocabulary signals that the value and error definitions belong to this average-reward perspective instead of being copied unchanged from the usual formulation.
One Sample Through an Approximation
Semi-gradient Sarsa becomes relevant when the action-value function is represented approximately rather than treated as a fully explicit table. A sampled transition is used to change the approximated action-value estimate, and the learning process changes the parameter vector that represents that approximation. The word semi-gradient identifies the algorithmic family; the important connection here is that Sarsa is being used with function approximation instead of only with a fully explicit table.
Following One Approximated Update
Generated example: Trace the role of one sampled transition when Sarsa uses an approximate action-value function.
Start with an approximation: The action-value function is represented through parameters rather than as a fully explicit table.
Use a sampled transition: The observed transition supplies learning information for the current approximated action-value estimate.
Change the parameters: Semi-gradient Sarsa uses the learning information to adjust the parameter vector that represents the approximation.
Evaluate the result: The result is a changed approximation, not the insertion of one independent table entry.
Generated example result: The essential flow is sampled transition, approximated action-value estimate, and parameter-vector update.
From One Step to n Steps
Sarsa and n-step Sarsa belong to the same broader family, but their backup targets differ in how far into the future the learning information reaches. The one-step view uses the immediate sampled transition and an estimate associated with the next step. An n-step view extends the backup across multiple steps, using rewards and estimates from farther into the future. In the average-reward setting, this parallel family includes differential versions of semi-gradient n-step Sarsa.
The n-step extension changes the reach of the backup, while the differential family changes the formulation used for average-reward learning. These are related but distinct ideas.
ε-Greedy Action Selection
Under ε-greedy action selection, learning chooses between an action currently estimated as best and exploratory actions. Therefore, the sequence of sampled transitions is shaped by both the current estimates and the exploratory part of the policy. The source reports the resulting convergence behavior for linear semi-gradient Sarsa with ε-greedy action selection: the parameters enter a bounded region near the best solution rather than converging in the usual sense.
What do you think happens?
Generated prediction: If linear semi-gradient Sarsa uses ε-greedy action selection, should you automatically describe its parameters as approaching one fixed value?
Reveal answer
Answer: No, the reported behavior is entry into a bounded region near the best solution.
The source explicitly distinguishes this behavior from convergence in the usual sense.
Bounded Region versus Fixed Point
Entering a bounded region near the best solution means that the parameter values remain within a neighborhood of that solution. Converging in the usual sense would mean approaching one limiting parameter value. These descriptions are not interchangeable: the first permits continued movement inside a bounded neighborhood, while the second describes approach to a single limit.
Treating a bounded region as identical to a single converged point.
A neighborhood can contain continuing variation, whereas usual convergence refers to approaching one limiting value.
Fix:
Describe the reported behavior as entering a bounded region near the best solution.Assuming average-reward learning reuses every ordinary value and error object unchanged.
The source identifies differential counterparts for each of these components.
Fix:
Use the differential terminology when discussing the average-reward formulation.Ignoring function representation when explaining semi-gradient Sarsa.
Semi-gradient Sarsa is connected in the source with approximate value-function representation.
Fix:
Explain that the sampled transition changes an approximation represented by parameters.
Practice Check
Generated practice: Explain, in your own words, why the phrase bounded region near the best solution should not be replaced with converges to the best solution when describing linear semi-gradient Sarsa with ε-greedy action selection.
Hints
- Contrast a neighborhood with one limiting value.
- Mention the parameter vector rather than only the action-value estimate.
- Use the reported behavior for the ε-greedy setting.
Generated practice: Organize these ideas into a learning pipeline: differential formulation, sampled transition, approximated action-value estimate, parameter-vector update, and bounded-region behavior.
Hints
- Start by identifying the average-reward formulation.
- Place the sampled transition before the approximation update.
- Treat bounded-region behavior as an interpretation of the longer-run parameter trajectory.
Key Takeaways
- Differential value functions are the average-reward counterparts of familiar value functions.
- Average-reward learning uses a parallel family of differential value functions, Bellman equations, and temporal-difference errors.
- Semi-gradient Sarsa is connected to function approximation because the action-value function is represented approximately and learning changes its parameter vector.
- n-step Sarsa extends the backup across multiple steps rather than restricting it to the one-step view.
- For linear semi-gradient Sarsa with ε-greedy action selection, the reported behavior is entry into a bounded region near the best solution, not convergence in the usual sense to one fixed point.
Key Takeaways
- Average-reward reinforcement learning requires differential counterparts of familiar value functions and temporal-difference components.
- Semi-gradient Sarsa applies the Sarsa idea when the action-value function is represented approximately through parameters.
- n-step Sarsa changes how far into the future the backup reaches, and differential n-step variants belong to the average-reward family.
- With linear semi-gradient Sarsa and ε-greedy action selection, parameters can enter a bounded region near the best solution without approaching one fixed limiting value.