Sarsa
Expected Sarsa is a Q-learning variation based on an expected next-state value.
One Next State, Three Backups
When an agent reaches a next state, several actions may be possible. Sarsa uses the value of the particular next action that was selected. Q-learning uses the highest-valued next action. Expected Sarsa takes a middle path: it considers all possible next actions and combines their values according to how likely they are under the relevant policy.
Expected Sarsa is a Q-learning variation based on an expected next-state value. It replaces Sarsa’s single sampled next action with an expected next-action value.
The central question is not merely which next state the agent reaches. The important question is how the algorithm decides what value that next state contributes to learning.
Tracing the Expected Backup
An Expected Sarsa update begins with the current state-action pair. The agent experiences a reward and reaches a next state. Once that next state is known, Expected Sarsa examines the actions available there. Instead of choosing only the action that happened to be sampled, it evaluates the possible next actions and weights each action-value estimate by that action’s likelihood under the current policy. The resulting expected next-action value is then used to update the earlier action-value estimate.
- Start from the current state-action pair.
- Observe the reward and the next state.
- List the possible actions in the next state.
- Use the relevant policy’s probability for each possible next action.
- Combine the next-action values according to those probabilities.
- Use the expected next-state value to improve the earlier action-value estimate.
Why Probabilities Matter
The current policy determines how much influence each possible next action has on the backup. An action with a high action-value estimate contributes strongly when the policy makes it likely, but an action with a lower value still contributes if the policy gives it some probability. Expected Sarsa therefore reflects the policy’s full distribution over next actions rather than representing the next state with one selected action or with only the best-valued action.
A policy-weighted next-state value
Suppose the next state has two possible actions. In this generated illustration, one action has value 8 and the policy makes it likely, while another action has value 4 and the policy makes it less likely.
Identify the action values: The possible next actions contribute values of 8 and 4.
Identify their policy probabilities: The generated illustration assigns probability 0.75 to the action valued at 8 and probability 0.25 to the action valued at 4.
Weight and combine: The weighted contributions are 6 and 1, so the expected next-action value is 7.
Interpret the result: The expected value is below 8 because the policy sometimes selects the action valued at 4. It is above 4 because the action valued at 8 is more likely.
The expected next-state value in this generated illustration is 7. Both actions matter, and the more probable action has the larger influence.
Sample, Average, or Maximum
| Method | Information from the next state | Role of the policy |
|---|---|---|
| Sarsa | The value of the particular next state-action pair selected by the agent | The selected action determines the backup |
| Expected Sarsa | The policy-weighted expected value of the possible next actions | The probabilities of the relevant policy determine each action’s influence |
| Q-learning | The maximum next action-value estimate | The target is greedy in the relationship described by the source |
Sarsa follows one realized next action. That makes its update depend on which action the agent happened to select. Expected Sarsa instead uses the possible next actions together, so its backup is policy-aware without reducing the next state to only its highest-valued action. Q-learning takes the greedy extreme by using the maximum next value.
Variance and Computation
A single sampled next action can make a Sarsa update depend heavily on that particular random choice. Expected Sarsa replaces that single-action backup with an average over the actions the policy might take. This removes the variance caused by random next-action selection from the update, although it requires additional computation to evaluate the possible actions and combine their values.
Policy Relationships
Expected Sarsa can be used on-policy or off-policy. It is on-policy when the policy that generates the agent’s behavior is also the target policy whose action probabilities supply the expected value. It is off-policy when one policy generates the experience while a different target policy supplies the action probabilities used for learning.
When the target policy is greedy, it assigns all probability to the action with the highest action-value estimate. The policy-weighted expectation then selects that maximum value, so this off-policy form of Expected Sarsa is exactly Q-learning.
Step Size in Practice
The step-size parameter controls how strongly a new update changes an existing action-value estimate. The reported cliff-walking results show an important contrast: for the deterministic state transitions described there, Expected Sarsa could safely use a step size of α = 1 without degradation of asymptotic performance. Sarsa performed well in the long run only at a small step size in that comparison, and that small value could make its short-term performance poor.
| Method | Reported step-size behavior in the cliff-walking comparison | Practical implication |
|---|---|---|
| Expected Sarsa | Could safely use α = 1 for the described deterministic transitions without asymptotic degradation | Less constrained by the step-size choice in the reported setting |
| Sarsa | Performed well in the long run only at a small step size | The small step size could make short-term performance poor |
Common Reasoning Errors
Treating Expected Sarsa as if it used only the next action that was selected.
That is the sample-based route associated with Sarsa. Expected Sarsa evaluates the possible next actions and combines their values according to policy likelihood.
Fix:
Ask which actions are possible in the next state and how much probability the relevant policy assigns to each one.Confusing a policy-weighted expectation with the maximum action value.
That describes the greedy maximum used by Q-learning, not the general Expected Sarsa backup.
Fix:
Retain every possible next action that has policy probability. Only when the target policy is greedy does the expectation become the maximum value.Assuming the behavior policy must always equal the target policy.
Expected Sarsa can be on-policy or off-policy, depending on whether the behavior and target policies are the same or different.
Fix:
Identify which policy generated the experience and which policy supplies the action probabilities for the expected value.Ignoring the computational cost of the expectation.
The variance reduction comes with additional computation because multiple possible next actions must be evaluated and combined.
Fix:
Compare both sides of the trade-off: more work per update and less variance from random next-action selection.
Check Your Understanding
An agent reaches a next state with three possible actions. The target policy assigns nonzero probability to all three. Which next-state information should Expected Sarsa use, and why would using only the action that happened to be selected change the nature of the update?
Hints
- Expected Sarsa considers the possible next actions rather than only one sampled action.
- The target policy’s probabilities determine how strongly each action-value contributes.
- Using only the selected action would make the update follow the sample-based route associated with Sarsa.
What do you think happens?
The target policy assigns all probability to the action with the highest next action-value. What does the Expected Sarsa next-state value become?
Reveal answer
Answer: The maximum next action-value.
When all target-policy probability is assigned to the highest-valued action, the policy-weighted expectation selects that value. In this greedy-target form, Expected Sarsa is exactly Q-learning.
Key Takeaways
- Expected Sarsa evaluates a next state with a policy-weighted expectation over possible next actions.
- Sarsa uses one sampled next action, while Q-learning uses the maximum next action value.
- The current target policy’s action probabilities determine how strongly each possible next action contributes.
- Averaging over possible next actions removes variance caused by the random selection of one next action, but requires additional computation.
- Expected Sarsa can be on-policy or off-policy, and its greedy-target off-policy form is exactly Q-learning.