ε-Greedy Policies
Sarsa control estimates the action values of its current behavior policy.
Learning Through Action Choice
An action-value estimate is useful only when the learner connects it to action selection. Sarsa control does this by estimating the action values of the policy it is currently using, then gradually changing that same policy toward greediness with respect to the estimates. The central idea is that acting, evaluating, and improving are linked rather than separated.
Sarsa control does not evaluate one fixed behavior policy while improving a different policy. It estimates the action values of its current behavior policy and adjusts that policy toward a greedy policy.
Sarsa's On-Policy Trace
Sarsa is called on-policy because the value estimate concerns the behavior policy itself. The policy selects the current action, receives a reward, selects the next action, and supplies that next action to the update. The behavior policy is not necessarily fixed: it can change toward greediness as learning proceeds. What makes the method on-policy is that the policy being evaluated is the policy being used to generate the experience.
Following One Sarsa Step
Trace the policy relationship when a learner uses its current policy to choose a current action and then a next action.
Choose the current action: The current behavior policy selects an action from the current situation.
Observe experience: The learner receives a reward as part of the observed transition.
Choose the next action: The same current policy selects the next action.
Update the estimate: The action-value estimate is updated using the next action selected by that policy.
The estimate concerns the policy that generated the actions, which is why the procedure is on-policy.
Moving Toward Greediness
An ε-greedy policy is defined relative to the current action-value estimates. The policy is therefore connected to Q-values: as the estimates change, the action preferences of the policy can change as well. In Sarsa control, the same policy is gradually changed toward greediness with respect to those estimates.
The source gives ε set to 1 divided by t as an example of a schedule that arranges for the policy to approach the greedy policy as learning progresses. Here, t represents the progression of learning. The important condition is not merely that the policy is called ε-greedy at one moment; it is that the policy becomes greedy in the limit while the required state-action exploration continues.
Conditions for Convergence
Sarsa's convergence result is conditional. The stated conditions are that every state-action pair is visited infinitely often and that the policy converges in the limit to the greedy policy. When those conditions hold, Sarsa converges with probability 1 to an optimal policy and an optimal action-value function.
| Condition | Why it matters | Consequence stated by the source |
|---|---|---|
| Every state-action pair is visited infinitely often | The learning process continues to receive experience for every pair | This is required by the convergence result |
| The policy becomes greedy in the limit | The policy approaches greediness with respect to its action-value estimates | This is required by the convergence result |
| Both conditions hold | The conditional convergence statement applies | Sarsa converges with probability 1 to an optimal policy and optimal action-value function |
The convergence claim depends on both exploration and limiting policy conditions.
Behavior and Target Policies
Off-policy n-step Q(σ) separates two roles. The behavior policy μ supplies the actions that generate experience. The target policy π determines the policy whose action values are being learned. This separation is the defining contrast with the on-policy Sarsa trace: the policy collecting experience and the policy being learned do not have to be the same.
In the stated algorithm, μ must assign positive probability to every state-action pair. The target policy π is initialized as ε-greedy with respect to Q, or it may be supplied as a fixed policy.
Importance Sampling and Mismatch
Experience collected under μ cannot simply be treated as though π had generated it. Importance sampling connects the two policies by correcting the update with ρ. A useful debugging path is to follow the observed action sequence, inspect the probabilities assigned by the behavior and target policies, track the ratio accumulation, and then inspect the resulting Q-value update.
Locating a Policy Mismatch
An observed action sequence was generated by μ, but the learner is updating Q-values for π. Identify where to look when the update differs from the expected result.
Check the generating policy: Confirm which actions in the sequence were supplied by μ.
Check the target policy: Determine what probability π assigns to each observed action.
Find the mismatch: At any action where the behavior and target policies assign different probabilities, the relationship between the observed experience and the target update changes.
Inspect ρ: Follow how the policy difference affects the importance-sampling correction and its accumulation.
Inspect the Q-value update: A changed correction can make the final update differ from an update that incorrectly treated μ-generated experience as direct π-generated experience.
The mismatch is not diagnosed by looking only at the final Q-value. Trace the action, both policy probabilities, ρ, and then the update.
Common Policy Mistakes
Calling Sarsa on-policy only because it uses action values.
The on-policy property comes from estimating the policy that generates the behavior, including the next action used by the update.
Fix:
Check whether the behavior policy and the policy being evaluated are the same.Treating the convergence result as unconditional.
The stated result requires every state-action pair to be visited infinitely often and the policy to become greedy in the limit.
Fix:
Verify both conditions before applying the convergence claim.Confusing μ with π in off-policy n-step Q(σ).
μ generates experience, while π is the target policy whose action values are learned.
Fix:
Label the experience source as μ and the learned policy as π.Ignoring importance sampling after observing a policy mismatch.
Importance sampling is the connection that corrects behavior-policy experience for target-policy learning.
Fix:
Trace the policy probabilities, ratio accumulation, and Q-value update.
Check Your Understanding
Explain, in your own words, why Sarsa is on-policy even though its policy changes during learning. Then describe the separate jobs of μ, π, and ρ in off-policy n-step Q(σ).
Hints
- For Sarsa, identify which policy selects the current and next actions.
- For off-policy learning, separate experience generation from the policy whose Q-values are learned.
- Mention that ρ corrects the update using the relationship between the two policies.
What do you think happens?
If a Sarsa policy is arranged so that ε becomes 1 divided by t as learning progresses, what limiting behavior is this intended to produce?
Reveal answer
Answer: It is an example of arranging for the policy to become greedy in the limit.
The source identifies ε set to 1 divided by t as an example schedule associated with the limiting policy condition.
Essential Takeaways
- Sarsa estimates the action values of its current behavior policy while gradually changing that policy toward greediness.
- Sarsa is on-policy because the behavior policy is also the policy being evaluated and used for the next-action part of the update.
- The stated convergence result requires every state-action pair to be visited infinitely often and the policy to become greedy in the limit.
- In off-policy n-step Q(σ), μ generates experience and π is the policy whose action values are learned.
- Importance sampling uses ρ to connect μ-generated experience to updates for π, so policy mismatches must be traced through the ratio and the Q-value update.
Key Takeaways
- Sarsa links action-value estimation with improvement of the same policy that generates behavior.
- Its on-policy character comes from using the behavior policy itself in the value estimate and next-action update.
- Convergence to an optimal policy and optimal action-value function is conditional on infinite visitation of every state-action pair and greediness in the limit.
- An ε-greedy policy can approach a greedy policy through a schedule such as ε equal to 1 divided by t.
- Off-policy n-step Q(σ) separates μ from π and uses importance sampling through ρ to correct behavior-policy experience for target-policy learning.