Concepts / Expected Sarsa

Expected Sarsa

Expected Sarsa replaces Sarsa's single sampled next action with an expected next-action value.

  • Programming

One Next Action or All Possibilities

When an agent reaches a next state, several actions may be possible. Sarsa bases its update on the one next action that the agent actually selects. Expected Sarsa takes a broader view: it considers the possible next actions and combines their action-value estimates according to how likely they are under a policy. This replaces a single sampled next action with an expected next-action value.

Expected Sarsa changes the question from “Which next action happened this time?” to “What value should the next state contribute when all policy-allowed next actions are considered?”

The Update Path

experience providesexperience reachesevaluate possible actionscombine using policy likelihoodsjoins target informationdrives updateCurrentstate-action pairQ estimateObserved rewardexperiencePolicy-weightedaction valuesall next actionsExpected next-actionvaluebackup informationUpdated Q-valueearlier pairNext statepossible actions
How does information move from the current state-action pair through the observed reward and next state into the expected target and updated Q-value?

The update begins with an earlier state-action pair and the experience obtained after taking that action: a reward and a next state. Once the next state is known, Expected Sarsa examines the actions available there. It does not use only the action sampled by the behavior process, and it does not keep only the action with the largest value. Instead, it combines the next action-values using their likelihood under the relevant policy. That combined value supplies the next-state information for updating the earlier estimate.

Sampled Versus Expected Backups

usesaverages according to policySarsaone sampled actionSelected next actionone Q-valueExpected Sarsapolicy-weighted actionsPossible next actionscombined Q-values
What is the difference between allowing one selected next action to determine the backup and combining all possible next actions under the policy?

Sarsa's backup depends on the particular next action selected in that experience. If the policy can select different actions, different samples can produce different updates even when the current state and earlier action are the same. Expected Sarsa replaces that single-action dependence with a policy-weighted combination of possible next actions. The random selection of one next action therefore no longer adds variance to the update.

Two Possible Next Actions

An agent reaches a next state with two possible actions. One action has value 8 and the other has value 4. The current policy makes the action valued at 8 more likely, but the action valued at 4 remains possible.

Sarsa's view: Sarsa uses the value of whichever one of the two actions was actually selected on this experience. A different sampled action can therefore change the update.

Expected Sarsa's view: Expected Sarsa includes both values. The value 8 contributes more because its action is more likely under the current policy, while the value 4 still contributes because its action remains possible.

Learning consequence: The expected backup reflects the policy's distribution over actions instead of letting one sample stand for the entire distribution.

Expected Sarsa uses a policy-weighted value for the next state; Sarsa uses the value of the one next action that was selected.

Why Probabilities Matter

paired withpaired withweights contributionweights contributionAction AQ-valuePolicy likelihood AweightExpected next-actionvaluecombined contributionAction BQ-valuePolicy likelihood Bweight
How do the current policy's action probabilities weight each next action-value to form the expected value?

The current policy's action probabilities are not decorative information. They determine how much each next action-value contributes to the expected value. A highly likely action has a larger influence, while an unlikely action has a smaller influence. An action with a lower value can still affect the backup if the policy assigns it some likelihood.

Relation to Q-Learning

uses target policyassigns target tousesExpected Sarsapolicy-weighted valuesGreedy target policyhighest-valued actionHighest next valueonly target contributionQ-learningmaximum next value
When the target policy always chooses the highest-valued action, how does the Expected Sarsa target become the Q-learning target?

Expected Sarsa can use different target policies. If the target policy is greedy, it assigns its target to the highest-valued next action. In that case, the policy-weighted expectation contains only the greedy action's contribution, so the Expected Sarsa target becomes the Q-learning target. This is why the source describes greedy-target off-policy Expected Sarsa as exactly Q-learning.

MethodNext-state information usedRole of action selection
SarsaValue of the particular next state-action pairUses the next action actually selected
Expected SarsaPolicy-weighted expectation over possible next actionsUses the relevant policy's action probabilities
Q-learningMaximum next action-valueUses a greedy target policy

On-Policy and Off-Policy Use

also supplies expectationexperience differs from targetBehavior policygenerates experienceTarget policysame policyBehavior policygenerates experienceTarget policydifferent policy
Which policy supplies the action probabilities used in the expectation, and how does that differ between on-policy and off-policy Expected Sarsa?

Expected Sarsa is on-policy when the policy generating the behavior is also the policy whose action probabilities form the expectation. It is off-policy when one policy generates the experience while a different target policy supplies the expected next-action value. This flexibility lets Expected Sarsa learn about a target policy that is different from the policy producing the observed behavior.

The phrase “expected” does not identify a single fixed policy relationship. The relevant question is which policy supplies the action probabilities used in the expectation.

Step Size and Variance

reported settingreported settingSarsasmall step sizeLong-run performanceperforms wellExpected Sarsaalpha equals 1Asymptoticperformanceno degradation
How does changing the step size alter the size and stability of Expected Sarsa and Sarsa updates when the next action is sampled versus averaged?

The step-size parameter controls how strongly an update changes an action-value estimate. In the reported deterministic cliff-walking setting, Expected Sarsa could safely use a step size of α = 1 without degradation of asymptotic performance. Sarsa, in contrast, performed well in the long run only at a small step size in that comparison, and the source reports that this small value could make Sarsa's short-term performance poor.

Expected Sarsa does pay for its lower selection variance with additional computation. Each update must calculate an expected next-action value rather than use only the sampled next action. The reported cliff-walking results show the trade-off: Expected Sarsa required more work per update but outperformed Sarsa and was less constrained by the step-size choice in that setting.

Mistakes in Identifying the Backup

  • Treating Expected Sarsa as if it used only the next action that happened to be selected

    That is the sample-based route associated with Sarsa. Expected Sarsa considers the possible next actions under a policy.

    Fix: Identify the relevant policy and combine the possible next action-values according to that policy's action probabilities.

  • Treating Expected Sarsa as identical to Q-learning in every configuration

    Expected Sarsa uses a policy-weighted expectation. It becomes Q-learning when the target policy is greedy, not in every policy configuration.

    Fix: Check whether the target policy is greedy before calling the backup a Q-learning backup.

  • Ignoring the action probabilities

    The expected value is shaped by how likely each action is under the relevant policy.

    Fix: Use the policy's action likelihoods to determine each action's contribution.

  • Assuming on-policy Expected Sarsa is the only form

    Expected Sarsa can also be off-policy, with one policy generating experience and another policy supplying the expectation.

    Fix: Compare the behavior policy with the target policy to classify the method.

Practice the Information Choice

MEDIUM

An agent has reached a next state with several available actions. The behavior process selects one action, but the current policy assigns different likelihoods to all of the available actions. Explain what information Sarsa uses, what information Expected Sarsa uses, and what information Q-learning uses for the update of the earlier state-action estimate.

Hints
  • Start by asking whether the method uses the sampled action, all possible actions, or only the highest-valued action.
  • For Expected Sarsa, identify the policy whose action probabilities form the expectation.
  • Remember that a greedy target policy makes the Expected Sarsa backup coincide with the Q-learning backup.

What do you think happens?

If two experiences begin with the same earlier state-action pair and reach the same next state, can Sarsa still produce different updates when different next actions are selected?

  • Yes
  • No
Reveal answer

Answer: Yes

Sarsa uses the value of the particular next action selected. A different sampled next action can therefore change the update. Expected Sarsa reduces this source of variation by using the policy-weighted expectation over possible next actions.

Summary

  1. Sarsa uses the value of the particular next action selected.
  2. Expected Sarsa uses the policy-weighted expected value of the possible next actions.
  3. Averaging over possible next actions removes variance caused by the random selection of one next action, at the cost of additional computation.
  4. Expected Sarsa is on-policy when behavior and target policies match, and off-policy when they differ.
  5. With a greedy target policy, Expected Sarsa becomes Q-learning because only the highest-valued next action contributes.
  6. In the reported deterministic cliff-walking comparison, Expected Sarsa outperformed Sarsa and was less constrained by the step-size choice.

Key Takeaways

  • Expected Sarsa replaces Sarsa's single sampled next action with a policy-weighted expectation over possible next actions.
  • The current policy's action probabilities determine how strongly each next action-value contributes.
  • This expectation reduces variance from random next-action selection but requires additional computation.
  • Expected Sarsa can be on-policy or off-policy, and its greedy-target form is exactly Q-learning.
  • In the reported cliff-walking setting, Expected Sarsa was less constrained by the step-size choice than Sarsa.