n-step TD prediction
n-step Sarsa is the n-step version of Sarsa and an on-policy TD control method.
From States to Actions
n-step TD prediction normally organizes its backup around states. When the same n-step idea is combined with Sarsa, the focus changes: the method works with state-action pairs. This produces n-step Sarsa, an on-policy temporal-difference control method.
The central construction change is simple but important: states are replaced by state-action pairs. The method still uses information from multiple steps, but it estimates action values rather than keeping the return in state-value form.
Tracing the Multi-step Backup
To understand the structure, imagine an agent moving through an alternating sequence of states and actions. The current state is paired with the action selected there. Over the next n steps, the backup gathers multi-step information. The quantity being updated is the current state-action pair, and the later part of the return is expressed with estimated action values.
An n-step Sarsa backup alternates states and actions and begins and ends with an action. This distinguishes it from a state-centered n-step backup, which begins and ends with a state.
Why Sarsa Changes the Return
Combining n-step prediction with Sarsa involves two linked changes. First, the backup is organized around state-action pairs. Second, the n-step return is redefined using estimated action values. The multi-step information is retained, but the estimated quantity is now tied to the action taken with the state.
A conceptual n-step Sarsa trace
Follow a generated sequence in which an agent encounters several alternating states and actions. Identify what is updated and what kind of value completes the n-step return.
1. Select the current pair: Start with the current state-action pair. This pair, rather than the state alone, is the object whose action value is updated.
2. Extend across n steps: Use information from multiple steps, including the rewards encountered along the sequence.
3. Preserve the action sequence: Keep the alternating state-action structure. The backup starts with an action and ends with an action.
4. Complete the return: Use an estimated action value associated with the later state-action pair instead of keeping the return in state-value form.
5. Improve control: Use an ε-greedy policy to select actions, allowing the n-step method to serve as an on-policy control method.
The result is n-step Sarsa: a multi-step, action-value method that updates state-action pairs and uses an ε-greedy policy for on-policy control.
Choosing Actions On Policy
n-step Sarsa uses an ε-greedy policy for action selection. This policy balances exploration and exploitation: it provides the policy component needed for control while keeping the method centered on the actions and action values in the Sarsa formulation.
The important point is the role of the policy, not a particular numerical action-selection rule. ε-greedy supplies the exploration-and-exploitation balance used to choose actions. Because n-step Sarsa is on-policy, the control method remains tied to the policy used for its action selection.
| Part of the method | What it contributes |
|---|---|
| n-step method | A multi-step backup structure |
| Sarsa | A focus on state-action pairs and action values |
| ε-greedy policy | A policy for balancing exploration and exploitation during action selection |
| Combined method | An on-policy temporal-difference control method called n-step Sarsa |
Mistakes in Reading the Backup
Treating n-step Sarsa as if it updated states rather than state-action pairs.
The main construction change is replacing states with state-action pairs.
Fix:
Identify the current pair as the object being updated and track the alternating sequence of states and actions.Keeping the n-step return in state-value form.
In n-step Sarsa, the return is redefined using estimated action values.
Fix:
Associate the later estimated value with a state-action pair.Assuming that the backup starts and ends with states.
n-step Sarsa backups alternate states and actions and begin and end with an action.
Fix:
Place the selected action at both ends of the backup structure.Leaving ε-greedy policy out of the control explanation.
The method uses an ε-greedy policy for action selection and is an on-policy control method.
Fix:
Explain how the policy balances exploration and exploitation while actions are selected.
When reading or drawing an n-step Sarsa backup, ask three questions in order: What state-action pair is being updated? Which later state-action pairs belong to the n-step sequence? Which estimated action value expresses the end of the return?
Check Your Understanding
Explain, in your own words, how n-step Sarsa differs from a state-centered n-step prediction method. Include the object being updated, the structure of the backup, the kind of value used in the return, and the role of the ε-greedy policy.
Hints
- Begin with the phrase state-action pair rather than state.
- Mention that the backup alternates states and actions and begins and ends with an action.
- Explain that the return uses estimated action values.
- State that ε-greedy balances exploration and exploitation and supplies the policy used for action selection.
What do you think happens?
A learner says, n-step Sarsa simply extends a state-value n-step backup across more steps. What essential detail is missing?
Reveal answer
Answer: The learner has missed the change from states to state-action pairs. n-step Sarsa also expresses the return using estimated action values and uses an ε-greedy policy for action selection.
The multi-step extension is only one part of the combination. Sarsa changes the estimated object and the backup structure, while ε-greedy supplies the policy component for on-policy control.
Essential Takeaways
- n-step Sarsa combines the multi-step structure of n-step methods with the Sarsa focus on actions.
- The method updates state-action pairs rather than states.
- Its backups alternate states and actions and begin and end with an action.
- The n-step return is expressed using estimated action values.
- An ε-greedy policy balances exploration and exploitation and makes n-step Sarsa an on-policy temporal-difference control method.
Key Takeaways
- n-step Sarsa is the n-step version of Sarsa and an on-policy temporal-difference control method.
- Its central structural change is replacing states with state-action pairs.
- The backup alternates states and actions and starts and ends with an action.
- The return uses estimated action values, while ε-greedy action selection balances exploration and exploitation.