Concepts / On-Policy and Off-Policy Control

On-Policy and Off-Policy Control

Sarsa learns action values one transition at a time.

  • Programming

Learning While Acting

Sarsa learns action values one transition at a time while an agent interacts with an environment. Instead of waiting until an episode finishes, it updates the value of the current state-action pair after each transition from a nonterminal state. This makes the order of events essential: a Sarsa update depends not only on the current state and action, but also on the next action selected by the policy.

The name Sarsa comes from the ordered event sequence S, A, R, S′, A′: current state, current action, reward, next state, and next action.

The Five-Event Update

A Sarsa transition begins with the current state and the action selected there. The agent acts, receives a reward, and observes the next state. It then selects the next action using its policy. Only after that next action has been selected does the update use the value of the resulting state-action pair. The five events are therefore connected, not interchangeable.

selectact and receiveobserveselect from policyuse next action valuecontributeScurrent stateAcurrent actionRrewardS′next stateA′next actionQ updatecurrent action value
How do the current state, current action, reward, next state, and next action connect to produce one Sarsa update?

Following One Sarsa Transition

Trace a generated transition in which an agent starts in state S, selects action A, receives reward R, reaches state S′, and then selects action A′.

Start: The current state-action pair is S and A. Its action value is the value that will be updated.

Interact: The agent takes A, receives R, and observes S′. These are the environmental consequences of the current action.

Select: From S′, the policy selects A′. Sarsa uses the value of this actual next action rather than an unrelated action.

Update: The reward and the value associated with S′ and A′ provide the information used to improve the value estimate for S and A.

Advance: The next state-action pair becomes the current pair for the following step.

One complete Sarsa update has used the ordered quintuple S, A, R, S′, A′, and the process can continue from the new state-action pair.

Prediction and Improvement

Sarsa applies temporal-difference prediction to control. The prediction part improves action-value estimates from experience for the current policy. The improvement part derives behavior from those estimates, so the policy can change locally as the estimates improve. These activities form a generalized policy iteration pattern: estimate values for the current behavior, then use those values to improve the behavior.

selects actionssupports TD updatesguideschanges behaviorCurrent policyExperienceAction valuesLocal improvement
How do value prediction and local policy improvement repeatedly interact during Sarsa control?

The two parts cannot be treated as separate stages that happen only once. The agent keeps interacting, updating values, and changing behavior. Exploration is important because the value estimates and policy improvements depend on the experience the agent gathers while acting.

Tracing a Mismatch

When a hand calculation or implementation disagrees with the expected Sarsa result, inspect the event sequence before checking only the final arithmetic. A mismatch may begin when the wrong current action is recorded, when the observed reward or next state is misread, or when the update uses an action other than the next action actually selected by the policy.

begin diagnosissequence differssequence matchesaction differsaction matchesverify updatecorrect event dataSarsa transitionCheck S A R S′ A′Check recorded valuesCheck next actionCheck TD targetCheck updated actionvalue
At which step does the calculated Sarsa target or updated action value differ from the expected result?

A Wrong Next Action

A generated calculation has the correct current state, current action, reward, and next state, but it uses the value of a different action at the next state.

Compare the event sequence: The current state, current action, reward, and next state agree with the observed transition.

Inspect action selection: The policy selected A′ at the next state, but the calculation uses the value of another action.

Locate the divergence: The calculation diverges when it substitutes a different next action, before the update is completed.

Repair the update: Use the value of the actual next action selected by the policy, then recompute the update.

The likely source of the disagreement is the event sequence, specifically the next action used in the Sarsa update, rather than the final arithmetic alone.

Policy Relationships

The on-policy and off-policy distinction concerns the relationship between action selection and learning. An on-policy method uses the current policy to select actions, and the learning loop remains tied to the policy whose returns are being predicted. Sarsa is the source material's example of on-policy TD control.

An off-policy method separates the policy being improved from the policy that selected the observed actions. This separation gives off-policy methods a way to address exploration without being restricted to the current policy. Q-learning and Expected Sarsa are classified as off-policy in the source material.

same policyseparate policiesCurrent policyBehavior policyCurrent policyImproved policy
What is the difference between the policy that generates behavior and the policy whose action values are learned?
MethodPolicy categoryRelationship emphasized by the source
SarsaOn-policyThe current policy selects actions.
Q-learningOff-policyThe policy being improved need not be the policy that selected the observed actions.
Expected SarsaOff-policyIt is classified as off-policy in the source material.

Policy categories identified in the source material.

Exploration During Control

An agent must choose actions while it is still learning. Its current policy determines how it acts, and the resulting experience supports value prediction. Those value estimates can then guide local policy improvement. The central design issue is exploration: the agent needs sufficient experience, but its action choices are also part of the behavior being evaluated or improved.

may selectmay selectcollectscollectssupports predictionsupports improvementCurrent policyExploratory actionGreedy actionNew experienceUpdated valuesImproved policy
How does choosing exploratory versus greedy actions affect the data collected and the policy being improved?

When analyzing a TD control method, ask two questions: which policy selects the actions that generate experience, and which policy's action values are being learned or improved? This distinction identifies the on-policy or off-policy category without confusing policy classification with the broader prediction-and-improvement process.

Practice Check

MEDIUM

A transition contains a current state, a current action, a reward, a next state, and a next action selected by the policy. Explain which event must be checked first if the calculated Sarsa result disagrees with the expected result, and classify Sarsa, Q-learning, and Expected Sarsa as on-policy or off-policy.

Hints
  • Start by writing the five events in their required order.
  • Check whether the update uses the actual next action selected by the policy.
  • Use the policy relationship, not the presence of prediction or improvement alone, to classify each method.

What do you think happens?

If the current state, current action, reward, and next state are all correct, but the update uses a different next action from the one selected by the policy, is the Sarsa calculation still using the correct event sequence?

  • Yes
  • No
Reveal answer

Answer: No

Sarsa uses the value of the actual next action selected by the policy. Substituting another action changes the update.

Key Takeaways

  1. Sarsa is an on-policy TD control method that learns action values one transition at a time.
  2. A Sarsa update follows the sequence S, A, R, S′, A′.
  3. The update uses the value of the actual next action selected by the policy.
  4. Sarsa combines TD value prediction with local policy improvement in a generalized policy iteration pattern.
  5. Exploration is central because the agent must gather experience while its values and policy are being improved.
  6. Sarsa is on-policy; Q-learning and Expected Sarsa are classified as off-policy in the source material.

Key Takeaways

  • Sarsa learns action values from each transition rather than waiting for an episode to finish.
  • Its defining event sequence is S, A, R, S′, A′.
  • TD control combines value prediction for current behavior with local policy improvement.
  • On-policy and off-policy control differ in the relationship between the policy selecting actions and the policy being learned or improved.
  • Exploration is the central design challenge because action selection produces the experience needed for learning.