On-policy and off-policy reinforcement learning
Expected Sarsa replaces Sarsa's single sampled next action with an expected next-action value.
One Next State, Different Targets
Sarsa, Q-learning, and Expected Sarsa all use experience to improve action-value estimates. Their important difference is how they extract learning information from the next state. Sarsa uses the next action that was actually selected. Expected Sarsa considers the expected value of the possible next actions under a policy. That distinction determines how much randomness enters each update and whether the method is on-policy or off-policy.
The central question is not only which next state the agent reaches. It is which next-action information the learning update uses after reaching that state.
Tracing the Next-Action Choice
Suppose an agent has reached a next state and its behavior policy can select more than one action there. Sarsa allows the action actually selected in that experience to determine the next-action contribution to the update. If a different action is selected on another experience, the contribution can change even when the next state and its estimated action values have not changed.
Expected Sarsa takes a broader view of that same next state. Instead of allowing one sampled action to represent the whole next state, it combines the values of the possible next actions according to their probabilities under the relevant policy. The update therefore uses an expected next-action value.
Averaging Away Selection Noise
The randomly selected next action can make Sarsa's target fluctuate. One experience may select an action with a relatively high estimated value, while another may select a different action with a lower estimated value. The update then inherits randomness from the action-selection event.
Expected Sarsa removes this particular source of variance by calculating the expected value of the possible next actions instead of using only the action sampled in one experience. This does not make the method free: calculating the expectation requires additional computation. The trade-off is extra work per update in exchange for removing the variance caused by random next-action selection.
One Next State, Two Update Routes
An agent reaches a next state where its policy gives more than one action a nonzero probability. Compare what Sarsa and Expected Sarsa use from this state.
Sarsa observes the selected action: The experience contains one action selected by the behavior policy. Sarsa uses that single action's estimated value for the next-action contribution.
Expected Sarsa considers the policy: Expected Sarsa considers the possible next actions and their policy probabilities, then combines their estimated values into an expected next-action value.
Compare the effect of randomness: If another experience selects a different action, Sarsa's next-action contribution can change because the sample changed. Expected Sarsa does not use that single sampled action as the entire representation of the next state.
Sarsa uses one sampled next action, while Expected Sarsa uses the policy-weighted expected value of possible next actions. Expected Sarsa therefore removes the variance introduced specifically by the random next-action selection, at additional computational cost.
Choosing the Policy Relationship
Expected Sarsa is on-policy when the policy used to generate the agent's behavior is also the target policy whose action values are being evaluated. In that case, the probabilities used to form the expected next-action value come from the same policy that generated the experience.
Expected Sarsa is off-policy when one policy generates the experience while a different target policy supplies the next-action probabilities used for learning. This gives Expected Sarsa flexibility: it can evaluate or improve the policy producing behavior, or it can learn about another target policy.
Step Size and Learning Trajectories
The step-size parameter controls how strongly each new experience changes an action-value estimate. The source reports an important difference in the deterministic state transitions of the cliff-walking setting. Expected Sarsa can safely use a step size of α = 1 without degradation of asymptotic performance, whereas Sarsa performs well in the long run only at a small step size in that comparison.
That small step size can make Sarsa's short-term performance poor. The reason is connected to the earlier distinction: Sarsa receives targets affected by the random sampled next action, so a large step size can make each noisy target strongly influence learning. Expected Sarsa pays more computation to average over possible next actions, and the reported results show that its performance is less constrained by the step-size choice in this setting.
| Method | Next-action information | Reported step-size behavior in deterministic cliff walking |
|---|---|---|
| Sarsa | Uses the action actually selected next | Performs well in the long run only at a small step size in the comparison; the small value can hurt short-term performance |
| Expected Sarsa | Uses the expected value of possible next actions under a policy | Can safely use α = 1 without degradation of asymptotic performance in the reported setting |
The comparison is specific to the deterministic cliff-walking results described by the source.
Common Misunderstandings
Treating Expected Sarsa as if it simply chooses a better single next action.
Expected Sarsa does not rely on one sampled next action. It combines the values of possible next actions according to their probabilities under the relevant policy.
Fix:
Describe the target as an expected next-action value, not as the value of one newly selected action.Assuming that removing selection variance makes Expected Sarsa free to compute.
The expectation requires additional computation.
Fix:
State the trade-off explicitly: Expected Sarsa does more work per update to remove variance from random next-action selection.Calling every use of Expected Sarsa on-policy.
The on-policy or off-policy classification depends on whether those policies are the same or different.
Fix:
Use on-policy when the behavior and target policies coincide; use off-policy when a different policy generates the experience.Saying that Expected Sarsa and Q-learning always produce the same target.
The exact Q-learning relationship applies when Expected Sarsa uses a greedy target policy.
Fix:
Limit the equivalence to the greedy-target off-policy form.Generalizing the cliff-walking step-size result to every task.
The source ties that result to deterministic state transitions in the reported cliff-walking setting.
Fix:
Treat the result as setting-specific evidence about reduced sensitivity, not as a universal guarantee.
Check Your Understanding
An agent reaches a next state. Its behavior policy can select several actions, but the learning algorithm is intended to evaluate a different target policy. Should the method be described as on-policy or off-policy Expected Sarsa, and what information should form the next-action contribution?
Hints
- Compare the policy generating the experience with the policy supplying the next-action probabilities.
- Remember that Expected Sarsa combines possible next-action values rather than using only the sampled action.
What do you think happens?
If Expected Sarsa uses a greedy target policy, does its expected next-action value remain different from the Q-learning target?
Reveal answer
Answer: No, because the greedy target selects the highest-valued next action.
With a greedy target policy, the expected next-action value reduces to the value of the highest-valued next action. The source identifies this greedy-target off-policy form of Expected Sarsa as exactly Q-learning.
Key Takeaways
- Sarsa uses the next action that was actually selected, while Expected Sarsa uses the expected value of possible next actions under a policy.
- Expected Sarsa removes variance caused by random next-action selection, but calculating the expectation requires additional computation.
- Expected Sarsa is on-policy when the behavior policy and target policy are the same, and off-policy when they are different.
- With a greedy target policy, Expected Sarsa becomes exactly Q-learning in the relationship described by the source.
- In the reported deterministic cliff-walking results, Expected Sarsa outperformed Sarsa and was less constrained by the step-size choice.
Key Takeaways
- Expected Sarsa replaces Sarsa's single sampled next action with an expected next-action value.
- The expectation reduces variance from random next-action selection at the cost of additional computation.
- Expected Sarsa can be on-policy or off-policy depending on whether its behavior and target policies are the same.
- When its target policy is greedy, Expected Sarsa is exactly Q-learning.
- The reported cliff-walking results show Expected Sarsa outperforming Sarsa and being less sensitive to the step-size choice.