Concepts / Sarsa

Sarsa

Expected Sarsa is a Q-learning variation based on an expected next-state value.

  • Programming

One Next State, Three Backups

When an agent reaches a next state, several actions may be possible. Sarsa uses the value of the particular next action that was selected. Q-learning uses the highest-valued next action. Expected Sarsa takes a middle path: it considers all possible next actions and combines their values according to how likely they are under the relevant policy.

Expected Sarsa is a Q-learning variation based on an expected next-state value. It replaces Sarsa’s single sampled next action with an expected next-action value.

Sarsasampled next actionExpected Sarsapolicy-weighted averageQ-learningmaximum next value
What information does each method use for the next action or next-state value?

The central question is not merely which next state the agent reaches. The important question is how the algorithm decides what value that next state contributes to learning.

Tracing the Expected Backup

An Expected Sarsa update begins with the current state-action pair. The agent experiences a reward and reaches a next state. Once that next state is known, Expected Sarsa examines the actions available there. Instead of choosing only the action that happened to be sampled, it evaluates the possible next actions and weights each action-value estimate by that action’s likelihood under the current policy. The resulting expected next-action value is then used to update the earlier action-value estimate.

  1. Start from the current state-action pair.
  2. Observe the reward and the next state.
  3. List the possible actions in the next state.
  4. Use the relevant policy’s probability for each possible next action.
  5. Combine the next-action values according to those probabilities.
  6. Use the expected next-state value to improve the earlier action-value estimate.
experiencereach next statecombine withweight valuesupdateCurrentstate-action pairexisting action-valueestimateReward and next statenew experiencePossible next actionsaction valuesPolicy probabilitieslikelihood of each actionExpected next-statevalueweighted combinationUpdated action-valueestimateimproved estimate
How does an Expected Sarsa update move from the current state-action pair through the reward and next-state action values to the updated estimate?

Why Probabilities Matter

The current policy determines how much influence each possible next action has on the backup. An action with a high action-value estimate contributes strongly when the policy makes it likely, but an action with a lower value still contributes if the policy gives it some probability. Expected Sarsa therefore reflects the policy’s full distribution over next actions rather than representing the next state with one selected action or with only the best-valued action.

A policy-weighted next-state value

Suppose the next state has two possible actions. In this generated illustration, one action has value 8 and the policy makes it likely, while another action has value 4 and the policy makes it less likely.

Identify the action values: The possible next actions contribute values of 8 and 4.

Identify their policy probabilities: The generated illustration assigns probability 0.75 to the action valued at 8 and probability 0.25 to the action valued at 4.

Weight and combine: The weighted contributions are 6 and 1, so the expected next-action value is 7.

Interpret the result: The expected value is below 8 because the policy sometimes selects the action valued at 4. It is above 4 because the action valued at 8 is more likely.

The expected next-state value in this generated illustration is 7. Both actions matter, and the more probable action has the larger influence.

valuevalueweighted contributionweighted contributionAction Avalue 8Probability 0.75policy weightExpected value 7weighted combinationAction Bvalue 4Probability 0.25policy weight
How are the current policy’s probabilities for possible next actions combined with their action values to produce the expected next-state value?

Sample, Average, or Maximum

MethodInformation from the next stateRole of the policy
SarsaThe value of the particular next state-action pair selected by the agentThe selected action determines the backup
Expected SarsaThe policy-weighted expected value of the possible next actionsThe probabilities of the relevant policy determine each action’s influence
Q-learningThe maximum next action-value estimateThe target is greedy in the relationship described by the source

Sarsa follows one realized next action. That makes its update depend on which action the agent happened to select. Expected Sarsa instead uses the possible next actions together, so its backup is policy-aware without reducing the next state to only its highest-valued action. Q-learning takes the greedy extreme by using the maximum next value.

selectaverage using probabilitieschoose maximumSarsaone sampled actionSelected action valueone next pairExpected Sarsaall possible actionsPolicy-weighted valueprobability-weightedQ-learninghighest-valued actionMaximum action valuegreedy value
What changes when the same next state is evaluated by Sarsa, Expected Sarsa, or Q-learning?

Variance and Computation

A single sampled next action can make a Sarsa update depend heavily on that particular random choice. Expected Sarsa replaces that single-action backup with an average over the actions the policy might take. This removes the variance caused by random next-action selection from the update, although it requires additional computation to evaluate the possible actions and combine their values.

select oneconsideraverageNext stateSarsa backupSampled actionone action valueNext stateExpected Sarsa backupPossible actionspolicy probabilitiesExpected valueaveraged action values
How does replacing one randomly selected next action with a probability-weighted average change the randomness of the target?

Policy Relationships

Expected Sarsa can be used on-policy or off-policy. It is on-policy when the policy that generates the agent’s behavior is also the target policy whose action probabilities supply the expected value. It is off-policy when one policy generates the experience while a different target policy supplies the action probabilities used for learning.

same policylearns aboutBehavior policygenerates experienceBehavior policygenerates experienceTarget policysame policyTarget policydifferent policy
How do the behavior policy that generates experience and the target policy that supplies action probabilities relate in each setting?

When the target policy is greedy, it assigns all probability to the action with the highest action-value estimate. The policy-weighted expectation then selects that maximum value, so this off-policy form of Expected Sarsa is exactly Q-learning.

concentrate probabilityexpected value equalsTarget policyprobabilities acrossactionsGreedy target policyall probability on bestactionMaximum next valueQ-learning target
Why does the Expected Sarsa target become the same as the Q-learning target when the target policy assigns all probability to the highest-valued action?

Step Size in Practice

The step-size parameter controls how strongly a new update changes an existing action-value estimate. The reported cliff-walking results show an important contrast: for the deterministic state transitions described there, Expected Sarsa could safely use a step size of α = 1 without degradation of asymptotic performance. Sarsa performed well in the long run only at a small step size in that comparison, and that small value could make its short-term performance poor.

MethodReported step-size behavior in the cliff-walking comparisonPractical implication
Expected SarsaCould safely use α = 1 for the described deterministic transitions without asymptotic degradationLess constrained by the step-size choice in the reported setting
SarsaPerformed well in the long run only at a small step sizeThe small step size could make short-term performance poor

Common Reasoning Errors

  • Treating Expected Sarsa as if it used only the next action that was selected.

    That is the sample-based route associated with Sarsa. Expected Sarsa evaluates the possible next actions and combines their values according to policy likelihood.

    Fix: Ask which actions are possible in the next state and how much probability the relevant policy assigns to each one.

  • Confusing a policy-weighted expectation with the maximum action value.

    That describes the greedy maximum used by Q-learning, not the general Expected Sarsa backup.

    Fix: Retain every possible next action that has policy probability. Only when the target policy is greedy does the expectation become the maximum value.

  • Assuming the behavior policy must always equal the target policy.

    Expected Sarsa can be on-policy or off-policy, depending on whether the behavior and target policies are the same or different.

    Fix: Identify which policy generated the experience and which policy supplies the action probabilities for the expected value.

  • Ignoring the computational cost of the expectation.

    The variance reduction comes with additional computation because multiple possible next actions must be evaluated and combined.

    Fix: Compare both sides of the trade-off: more work per update and less variance from random next-action selection.

Check Your Understanding

MEDIUM

An agent reaches a next state with three possible actions. The target policy assigns nonzero probability to all three. Which next-state information should Expected Sarsa use, and why would using only the action that happened to be selected change the nature of the update?

Hints
  • Expected Sarsa considers the possible next actions rather than only one sampled action.
  • The target policy’s probabilities determine how strongly each action-value contributes.
  • Using only the selected action would make the update follow the sample-based route associated with Sarsa.

What do you think happens?

The target policy assigns all probability to the action with the highest next action-value. What does the Expected Sarsa next-state value become?

  • The value of a randomly selected action
  • The probability-weighted average of several actions with nonzero probability
  • The maximum next action-value
Reveal answer

Answer: The maximum next action-value.

When all target-policy probability is assigned to the highest-valued action, the policy-weighted expectation selects that value. In this greedy-target form, Expected Sarsa is exactly Q-learning.

Key Takeaways

  • Expected Sarsa evaluates a next state with a policy-weighted expectation over possible next actions.
  • Sarsa uses one sampled next action, while Q-learning uses the maximum next action value.
  • The current target policy’s action probabilities determine how strongly each possible next action contributes.
  • Averaging over possible next actions removes variance caused by the random selection of one next action, but requires additional computation.
  • Expected Sarsa can be on-policy or off-policy, and its greedy-target off-policy form is exactly Q-learning.