Concepts / Sarsa and On-Policy TD Control

Sarsa and On-Policy TD Control

TD control joins value prediction with local policy improvement.

  • Programming

From Prediction to Control

A TD method can learn from experience, but control requires more than estimating how good a situation is. The learner must also improve the policy that chooses actions. This creates a central design question: while the learner explores to gather experience, which policy is it actually learning to evaluate and improve?

TD control joins two activities. Value prediction works toward accurate return predictions for the current policy. Policy improvement uses the current value function to make action choices better. Because both activities rely on experience, exploration cannot be treated as an unrelated detail. The method must account for how actions are explored while learning and improving.

learn frominformguiderefreshExperienceobservations and rewardsValue predictionreturn estimatesPolicy improvementbetter action choicesNew experienceupdated decisions
How do value prediction and local policy improvement support one another over repeated experience?

The word control signals the addition of policy improvement. A method that only predicts values for a fixed policy is not yet performing the full TD control process described here.

Following One Experience

A useful way to understand Sarsa is to follow one experienced transition conceptually. The learner starts with a state, selects an action, observes a reward and a next state, and then considers the next action selected by the policy being used for behavior. That sequence connects the experience to the value estimate being learned.

chooseproduceaccompanychoose nextStatecurrent situationActionselected actionRewardobserved outcomeNext stateresulting situationNext actionselected under the policy
How do the current state, selected action, outcome, and next selected action form the experience used by an on-policy method?

The important on-policy idea is not merely that the learner observes an action. The action choices used to generate experience are part of the same policy arrangement that the learner evaluates and improves. Therefore, exploration remains inside the policy being learned rather than being assigned to a completely separate behavior role.

Exploration Changes the Learning Problem

Exploration means that the learner does not simply exploit its current knowledge on every decision. It gathers experience while trying actions whose value may still be uncertain. That experience influences the value predictions, while the resulting predictions influence later policy improvement. The same learning process therefore has to manage both information gathering and better decision making.

selectscreatesinformssupportsExploring policyexploration andexploitationActionselected choiceExperiencereward and next situationValue estimatesupdated predictionsImproved policyrevised action choices
How does an exploratory choice affect the resulting experience and the policy that is being learned?

A policy must learn while it explores

Imagine an agent choosing between two actions while its value estimates are still incomplete. How should you interpret an experience produced by an exploratory choice?

Choose: The current policy selects an action while balancing exploration with exploitation.

Observe: The learner receives experience from that choice, including what happened afterward.

Predict: The experience contributes to return predictions for the policy arrangement being evaluated.

Improve: The current value information is used to improve future action choices.

Exploration is part of the control design because it determines which experience is gathered and how the policy being learned is connected to that experience.

Two Policy Roles

On-policy TD control uses the same policy for exploration and exploitation. The learner evaluates and improves behavior within that single policy arrangement.

Off-policy TD control uses different policies for exploration and exploitation. One policy is associated with generating experience, while another is associated with exploiting what has been learned.

same policy roledifferent policy rolesOne policyexploration andexploitationExploring policygenerates experienceOne policyevaluated and improvedExploiting policyuses what was learned
What differs between the policy that generates experience and the policy whose values are being learned?
QuestionOn-policy controlOff-policy control
How many policy roles are distinguished?One policy arrangementSeparate exploration and exploitation policies
What generates experience?The same policy arrangement being evaluated and improvedThe policy associated with exploration
What is the defining test?Exploration and exploitation use the same policyExploration and exploitation use different policies

Do not ask whether a method explores at all. Both policy arrangements must account for exploration. Ask instead whether the policy used for exploration is the same as the policy used for exploitation.

Classifying the Main Methods

The policy-role test places the three methods in two groups. Sarsa is on-policy because exploration and exploitation remain within one policy arrangement. Q-learning and Expected Sarsa are off-policy as presented in this material because exploration and exploitation are assigned different policy roles.

classified byclassified byclassified bySarsaon-policySame policyexploration andexploitationQ-learningoff-policyDifferent policiesexploration andexploitationExpected Sarsaoff-policy
Which policy relationship determines the classification of each method?
MethodClassificationPolicy relationship
SarsaOn-policyExploration and exploitation use the same policy arrangement
Q-learningOff-policyExploration and exploitation use different policy roles
Expected SarsaOff-policyExploration and exploitation use different policy roles

These classifications follow the policy relationship stated in the source material.

Reading the Control Process

Consider an agent whose current value information favors one action but is not yet complete. The agent follows its exploration policy, receives experience, and uses that experience to improve its predictions. The updated predictions then support a better policy. In an on-policy arrangement, the policy producing the next experience is also the policy arrangement being evaluated and improved. In an off-policy arrangement, the experience-generating role and the exploitation role are kept distinct.

updateguidegeneraterefineExperiencecurrent roundValue estimatesprediction informationAction choicespolicy improvementExperiencenext round
How does each round of value estimation provide information for policy improvement and later experience?

When classifying a TD control method, first identify the policy that generates experience and the policy associated with exploitation. Only after identifying those roles should you assign the method to the on-policy or off-policy group.

EASY

For each method below, classify it as on-policy or off-policy and explain the policy relationship that justifies your answer: Sarsa, Q-learning, and Expected Sarsa.

Hints
  • Use the same-policy versus different-policies test.
  • Do not classify a method from its name or perceived implementation complexity.
  • Remember that the source identifies Sarsa as on-policy and Q-learning and Expected Sarsa as off-policy.

Common Classification Errors

  • Treating TD control as value prediction only

    TD control joins value prediction with local policy improvement.

    Fix: Explain both parts: prediction improves return estimates, and policy improvement uses current value information to improve action choices.

  • Assuming exploration disappears in off-policy control

    Off-policy control separates exploration and exploitation; it does not remove exploration.

    Fix: Identify the exploring policy and the separate exploitation policy.

  • Classifying algorithms from their names

    The classification depends on the relationship between the exploration policy and the exploitation policy.

    Fix: Apply the policy-role test before assigning a category.

  • Putting Q-learning and Expected Sarsa in the on-policy group

    The material identifies both Q-learning and Expected Sarsa as off-policy.

    Fix: Treat them as off-policy because exploration and exploitation have different policy roles in this presentation.

  • Treating actor-critic as another name for off-policy control

    Actor-critic is presented as a related third approach that combines an actor and a critic.

    Fix: Keep actor-critic separate from the on-policy and off-policy classifications covered here.

Summary

  1. TD control combines value prediction with local policy improvement.
  2. Exploration matters because experience is used both to improve predictions and to guide better action choices.
  3. On-policy methods use the same policy for exploration and exploitation; Sarsa is the identified example.
  4. Off-policy methods use different policy roles for exploration and exploitation; Q-learning and Expected Sarsa are placed in this group.
  5. The reliable classification method is to compare the policy generating experience with the policy associated with exploitation.

Key Takeaways

  • TD control is the combination of value prediction and policy improvement.
  • Exploration must be considered because it determines how experience is gathered while the learner improves its decisions.
  • Sarsa is on-policy because exploration and exploitation use the same policy arrangement.
  • Q-learning and Expected Sarsa are off-policy as presented because exploration and exploitation use different policy roles.
  • Classify a method by comparing its experience-generating policy with its exploitation policy.