Sarsa algorithm
Calculate the n-Step TD Error before applying the semi-gradient Sarsa update.
The Two-Stage Idea
Sarsa is best understood by separating one learning step into two connected stages. First, the algorithm calculates the n-Step TD Error. Then, it applies the semi-gradient Sarsa update using the result of that calculation. The stages are connected, but they are not the same operation.
The essential order is: calculate the n-Step TD Error first, then apply the semi-gradient Sarsa update.
Error Before Update
The n-Step TD Error prepares the learning signal for the next part of the algorithm. Until that calculation is complete, the semi-gradient Sarsa update has not yet been applied for this step. Once the error has been calculated, the procedure is ready to begin the update stage.
Tracing One Learning Step
From experience to parameter update
Trace the order of operations in one high-level Sarsa learning step.
Gather the learning information: The procedure begins with experience gathered by the policy being improved. In a high-level trace, this experience supplies the information used by the learning procedure.
Calculate the n-Step TD Error: The algorithm calculates the n-Step TD Error. This completes the stage that prepares the learning signal.
Apply the semi-gradient Sarsa update: After the error is available, the algorithm applies the semi-gradient Sarsa update. This is the parameter-updating stage.
The n-Step TD Error calculation comes before the semi-gradient Sarsa update because the calculated error supplies the point at which the update begins.
This trace does not assign numerical values to the error or the parameters. Its purpose is to make the control flow explicit: information leads to error calculation, and the calculated error leads to the update.
Information Flow
The diagram represents the high-level dependency described by the source. States, actions, rewards, and the n-step estimate are treated as information involved in the error-calculation stage. The resulting n-Step TD Error then feeds the semi-gradient Sarsa update. The important guaranteed relationship is the final one: error calculation precedes parameter updating.
Sarsa as On-Policy Control
Sarsa is an on-policy temporal-difference control method. On-policy means that learning uses experiences gathered by the same policy that is being improved.
The word on-policy describes the relationship between the policy producing the learning experiences and the policy being improved. In Sarsa, those are the same policy. This classification is determined by examining the source of the experiences, not merely by memorizing the algorithm's name.
On-Policy and Off-Policy
| Category | Relationship | Examples from the source |
|---|---|---|
| On-policy TD control | The experiences are gathered by the same policy being improved. | Sarsa |
| Off-policy TD control | The experiences are gathered without following the same policy being improved. | Q-learning and Expected Sarsa |
The Wider TD Family
Temporal-difference control methods are used to control an agent in reinforcement learning. They connect two processes: one drives the value function toward accurately predicting returns for the current policy, while another improves the policy locally with respect to that value function. Sarsa belongs to this broader family and is specifically on-policy.
The source also identifies actor-critic methods as another extension of temporal-difference methods. This places actor-critic methods alongside, rather than inside the definition of, the specific Sarsa procedure discussed here.
Common Mistakes
Treating the n-Step TD Error calculation and the semi-gradient Sarsa update as one indistinguishable operation.
The source distinguishes two connected stages and gives them a definite order.
Fix:
Name the stages separately: calculate the n-Step TD Error, then apply the semi-gradient Sarsa update.Applying the semi-gradient Sarsa update before calculating the n-Step TD Error.
The calculated error supplies the input or starting point for the following update stage.
Fix:
Complete the error calculation first, then begin the update.Calling Sarsa off-policy simply because it is a TD control method.
TD control includes both on-policy and off-policy approaches.
Fix:
Check the relationship between the experience-generating policy and the policy being improved. Sarsa is on-policy.Defining on-policy by the algorithm's name instead of by the source of experience.
The important test is whether experiences come from the same policy being improved.
Fix:
Use the policy-relationship test whenever you classify a TD control method.
Check Your Understanding
A learner says: The semi-gradient Sarsa update should happen first, because the update is the main learning operation. Correct the learner's sequence and explain the role of the n-Step TD Error.
Hints
- Identify the two distinct stages.
- Ask which stage prepares the learning signal for the other stage.
- State the order explicitly.
A method gathers experiences using one policy but improves a different policy. Is this relationship on-policy or off-policy? Then contrast the answer with Sarsa.
Hints
- Compare the experience-generating policy with the policy being improved.
- Use Sarsa as the on-policy reference point.
Key Takeaways
- The n-Step TD Error and the semi-gradient Sarsa update are two connected but distinct stages.
- The n-Step TD Error is calculated first because it prepares the learning signal for the following update.
- Sarsa is an on-policy temporal-difference control method: the policy generating experience is the same policy being improved.
- Off-policy TD control uses experience gathered without following the same policy being improved; Q-learning and Expected Sarsa are examples identified in the source.
- Actor-critic methods are another extension of temporal-difference methods.
Key Takeaways
- Calculate the n-Step TD Error before applying the semi-gradient Sarsa update.
- The error calculation prepares the learning signal that the update stage uses.
- Sarsa is on-policy because it learns from experiences gathered by the policy being improved.
- Off-policy methods learn from experiences not gathered by the same policy being improved.
- Sarsa is part of the broader TD control family, while actor-critic methods are another TD extension.