On-Policy Reinforcement Learning
Sarsa control estimates the action values of its current behavior policy.
Learning Through the Acting Policy
On-policy reinforcement learning links two activities that are often discussed separately: estimating how good actions are and improving the policy that chooses those actions. In Sarsa control, the learner estimates the action values of the behavior policy it is currently using. That same policy is then gradually changed toward greediness with respect to the action-value estimates.
The central idea is that acting and improving are linked. Sarsa does not evaluate one policy while following a separately chosen policy; it evaluates the policy that is producing its behavior.
The Sarsa Learning Cycle
The cycle begins with the behavior policy currently used by the learner. That policy selects actions, and the resulting experience contributes to the learner's action-value estimates. Those estimates influence a policy adjustment: the policy is gradually moved toward choosing actions greedily with respect to the estimates. The adjusted policy then produces later experience.
Why Sarsa Is On-Policy
On-policy means that the action-value estimate concerns the behavior policy itself. In Sarsa control, the learner estimates the values of the policy it is currently using, rather than estimating a separately selected policy.
The policy is not fixed throughout learning. It begins as a behavior policy, supplies the learner's experience, and is gradually changed toward a greedy policy with respect to the current action-value estimates. Therefore, on-policy describes both what is being evaluated and how that evaluated policy changes over time.
Evaluation and Improvement Together
A Policy That Changes as Estimates Improve
Suppose the learner's current action-value estimates make one action appear better than another in a particular situation. How does Sarsa use that information?
Estimate the current behavior: The learner uses experience from its current behavior policy to estimate action values. The estimate concerns the policy that is actually being used.
Compare available actions: The current estimates indicate which action is presently preferred from the value-estimation perspective.
Adjust the policy: The behavior policy is gradually changed toward greediness with respect to the action-value estimates. The preferred action therefore becomes more influential in later policy choices.
Continue learning: The adjusted policy generates later experience, and that later experience is used to keep estimating and improving the same changing behavior policy.
Sarsa combines policy evaluation and policy improvement in one continuing loop: the policy supplies experience, its values are estimated, and those estimates guide changes to that policy.
This example illustrates the defining relationship without requiring the policy to become greedy immediately. Sarsa gradually changes the policy, so the policy being evaluated can develop over the course of learning.
Moving Epsilon Toward Greediness
An ε-greedy policy balances greedy action selection with exploration. The source gives ε = 1/t as an example of a schedule that makes the policy approach greediness as learning progresses. As t increases, ε becomes smaller, so the policy moves toward the greedy policy in the limit.
Convergence Requirements
Sarsa's convergence result is conditional. It is not an unconditional promise for every possible policy schedule. The stated conditions are that every state-action pair is visited infinitely often and that the policy becomes greedy in the limit.
- Every state-action pair must be visited infinitely often.
- The policy must converge in the limit to the greedy policy.
- Under these conditions, Sarsa converges with probability 1 to an optimal policy and an optimal action-value function.
A policy can be moving toward greediness without satisfying the convergence statement if some state-action pairs are not visited infinitely often. The convergence claim depends on both continued coverage of every state-action pair and limiting greediness.
Common Reasoning Errors
Treating Sarsa as if it evaluates a fixed policy.
Sarsa control continually estimates the current behavior policy while gradually changing that policy toward greediness.
Fix:
Remember that policy evaluation and policy improvement are linked in Sarsa control.Calling Sarsa on-policy merely because it uses action values.
The important distinction is which policy the action values concern.
Fix:
Sarsa is on-policy because its action-value estimates concern the behavior policy that is currently being used.Assuming that ε-greedy means permanently random behavior.
The source gives ε = 1/t as an example in which ε decreases as learning progresses and the policy approaches greediness.
Fix:
Think of ε as adjustable; with the stated schedule, the policy becomes greedy in the limit.Treating convergence as guaranteed without conditions.
The convergence result requires every state-action pair to be visited infinitely often and the policy to become greedy in the limit.
Fix:
State the conditions whenever you state the convergence result.
Check Your Understanding
Explain in your own words why Sarsa is called an on-policy control algorithm. Then describe what happens to an ε-greedy policy when ε is set to 1/t and t increases.
Hints
- Identify which policy supplies the experience used for the action-value estimate.
- Distinguish the policy being evaluated from a separately selected policy.
- Consider what happens to ε as t becomes larger.
A learner's policy is becoming greedier, but one state-action pair is visited only a finite number of times. Does the stated Sarsa convergence result apply? Explain why or why not.
Hints
- Recall the requirement concerning every state-action pair.
- Separate the policy's limiting behavior from the visitation requirement.
Key Takeaways
- Sarsa control estimates the action values of the behavior policy it is currently using. Those estimates guide a gradual change of the same policy toward greediness, which is why Sarsa is on-policy. An ε-greedy policy can approach a greedy policy by using ε = 1/t as learning progresses. The convergence result is conditional: every state-action pair must be visited infinitely often, and the policy must become greedy in the limit. When those conditions hold, Sarsa converges with probability 1 to an optimal policy and optimal action-value function.
Key Takeaways
- Sarsa estimates the action values of the behavior policy it is currently using.
- The same policy is gradually improved toward greediness with respect to its action-value estimates.
- Sarsa is on-policy because the policy being evaluated is the policy generating behavior.
- An ε-greedy policy can approach a greedy policy by setting ε to 1/t as learning progresses.
- With infinite visits to every state-action pair and a policy that becomes greedy in the limit, Sarsa converges with probability 1 to an optimal policy and optimal action-value function.