On-Policy and Off-Policy Control
Sarsa learns action values one transition at a time.
Learning While Acting
Sarsa learns action values one transition at a time while an agent interacts with an environment. Instead of waiting until an episode finishes, it updates the value of the current state-action pair after each transition from a nonterminal state. This makes the order of events essential: a Sarsa update depends not only on the current state and action, but also on the next action selected by the policy.
The name Sarsa comes from the ordered event sequence S, A, R, S′, A′: current state, current action, reward, next state, and next action.
The Five-Event Update
A Sarsa transition begins with the current state and the action selected there. The agent acts, receives a reward, and observes the next state. It then selects the next action using its policy. Only after that next action has been selected does the update use the value of the resulting state-action pair. The five events are therefore connected, not interchangeable.
Following One Sarsa Transition
Trace a generated transition in which an agent starts in state S, selects action A, receives reward R, reaches state S′, and then selects action A′.
Start: The current state-action pair is S and A. Its action value is the value that will be updated.
Interact: The agent takes A, receives R, and observes S′. These are the environmental consequences of the current action.
Select: From S′, the policy selects A′. Sarsa uses the value of this actual next action rather than an unrelated action.
Update: The reward and the value associated with S′ and A′ provide the information used to improve the value estimate for S and A.
Advance: The next state-action pair becomes the current pair for the following step.
One complete Sarsa update has used the ordered quintuple S, A, R, S′, A′, and the process can continue from the new state-action pair.
Prediction and Improvement
Sarsa applies temporal-difference prediction to control. The prediction part improves action-value estimates from experience for the current policy. The improvement part derives behavior from those estimates, so the policy can change locally as the estimates improve. These activities form a generalized policy iteration pattern: estimate values for the current behavior, then use those values to improve the behavior.
The two parts cannot be treated as separate stages that happen only once. The agent keeps interacting, updating values, and changing behavior. Exploration is important because the value estimates and policy improvements depend on the experience the agent gathers while acting.
Tracing a Mismatch
When a hand calculation or implementation disagrees with the expected Sarsa result, inspect the event sequence before checking only the final arithmetic. A mismatch may begin when the wrong current action is recorded, when the observed reward or next state is misread, or when the update uses an action other than the next action actually selected by the policy.
A Wrong Next Action
A generated calculation has the correct current state, current action, reward, and next state, but it uses the value of a different action at the next state.
Compare the event sequence: The current state, current action, reward, and next state agree with the observed transition.
Inspect action selection: The policy selected A′ at the next state, but the calculation uses the value of another action.
Locate the divergence: The calculation diverges when it substitutes a different next action, before the update is completed.
Repair the update: Use the value of the actual next action selected by the policy, then recompute the update.
The likely source of the disagreement is the event sequence, specifically the next action used in the Sarsa update, rather than the final arithmetic alone.
Policy Relationships
The on-policy and off-policy distinction concerns the relationship between action selection and learning. An on-policy method uses the current policy to select actions, and the learning loop remains tied to the policy whose returns are being predicted. Sarsa is the source material's example of on-policy TD control.
An off-policy method separates the policy being improved from the policy that selected the observed actions. This separation gives off-policy methods a way to address exploration without being restricted to the current policy. Q-learning and Expected Sarsa are classified as off-policy in the source material.
| Method | Policy category | Relationship emphasized by the source |
|---|---|---|
| Sarsa | On-policy | The current policy selects actions. |
| Q-learning | Off-policy | The policy being improved need not be the policy that selected the observed actions. |
| Expected Sarsa | Off-policy | It is classified as off-policy in the source material. |
Policy categories identified in the source material.
Exploration During Control
An agent must choose actions while it is still learning. Its current policy determines how it acts, and the resulting experience supports value prediction. Those value estimates can then guide local policy improvement. The central design issue is exploration: the agent needs sufficient experience, but its action choices are also part of the behavior being evaluated or improved.
When analyzing a TD control method, ask two questions: which policy selects the actions that generate experience, and which policy's action values are being learned or improved? This distinction identifies the on-policy or off-policy category without confusing policy classification with the broader prediction-and-improvement process.
Practice Check
A transition contains a current state, a current action, a reward, a next state, and a next action selected by the policy. Explain which event must be checked first if the calculated Sarsa result disagrees with the expected result, and classify Sarsa, Q-learning, and Expected Sarsa as on-policy or off-policy.
Hints
- Start by writing the five events in their required order.
- Check whether the update uses the actual next action selected by the policy.
- Use the policy relationship, not the presence of prediction or improvement alone, to classify each method.
What do you think happens?
If the current state, current action, reward, and next state are all correct, but the update uses a different next action from the one selected by the policy, is the Sarsa calculation still using the correct event sequence?
Reveal answer
Answer: No
Sarsa uses the value of the actual next action selected by the policy. Substituting another action changes the update.
Key Takeaways
- Sarsa is an on-policy TD control method that learns action values one transition at a time.
- A Sarsa update follows the sequence S, A, R, S′, A′.
- The update uses the value of the actual next action selected by the policy.
- Sarsa combines TD value prediction with local policy improvement in a generalized policy iteration pattern.
- Exploration is central because the agent must gather experience while its values and policy are being improved.
- Sarsa is on-policy; Q-learning and Expected Sarsa are classified as off-policy in the source material.
Key Takeaways
- Sarsa learns action values from each transition rather than waiting for an episode to finish.
- Its defining event sequence is S, A, R, S′, A′.
- TD control combines value prediction for current behavior with local policy improvement.
- On-policy and off-policy control differ in the relationship between the policy selecting actions and the policy being learned or improved.
- Exploration is the central design challenge because action selection produces the experience needed for learning.