n-Step Methods for Reinforcement Learning
n-step semi-gradient Sarsa joins a semi-gradient Sarsa control method with n-step bootstrapping.
Why Look Beyond One Step
An estimate does not have to wait for every relevant piece of information before it can be revised. n-step semi-gradient Sarsa combines a semi-gradient Sarsa control method with n-step bootstrapping. The value of n determines how much subsequent estimate information contributes to learning. This creates a choice between relying on a shorter sequence of experience and allowing information from farther ahead to influence the current estimate.
The important design question is not simply whether to bootstrap. It is how far later experience should be allowed to contribute before the current estimate is adjusted.
Following an n-Step Update
A single n-step update can be understood as a short information trail. The agent begins with a state and action. It then experiences subsequent rewards and states. The chosen value of n determines how much of this later experience is considered before the estimate for the earlier state-action situation is revised. The update therefore combines information sampled from the task with a later estimate rather than waiting for all relevant information to arrive independently.
Tracing a Four-Step Update
An agent is updating the estimate for an earlier state-action pair and uses n = 4. Trace what information participates in the update.
Start at the earlier state-action pair: This is the estimate that will be revised.
Collect subsequent experience: The agent follows the later sequence of states and rewards relevant to the four-step look-ahead.
Use the later estimate: The update includes subsequent estimate information rather than relying only on the original estimate.
Revise the earlier estimate: The information gathered over the chosen horizon is used to adjust the estimate for the earlier state-action pair.
With n = 4, four-step look-ahead determines how much later experience contributes before the earlier estimate is adjusted. The example illustrates the role of n without assigning numerical rewards or values.
How Bootstrapping Moves Information
Bootstrapping means that an estimate is updated using subsequent estimates. A later state is therefore more than another observation from the environment: its estimate becomes part of the information used to revise the current estimate. In this way, information can move backward through the sequence of estimates. The value of n controls how far this process looks before the update is formed.
Changing the Horizon with n
Changing n changes the balance between sampled rewards and later estimates. A smaller n forms the update after a shorter look ahead, so the update depends more quickly on a nearby later estimate. A larger n allows a longer sequence of subsequent experience to contribute before the update is formed. The choice is not simply between n = 1 and using as many steps as possible; the reported Mountain Car results show that an intermediate choice can be better than either extreme.
| Choice of n | Look-ahead | Interpretation |
|---|---|---|
| n = 1 | Shorter | The update is formed with a short amount of subsequent information. |
| n = 4 | Intermediate | The update uses an intermediate amount of later experience and bootstrapping. |
| n = 8 | Longer | The update allows more subsequent experience to contribute before adjustment. |
Mountain Car Evidence
The reported Mountain Car comparison favors multi-step learning over n = 1. In that comparison, n = 8 learned faster and reached better asymptotic performance than n = 1. A more detailed study found n = 4 to be the best intermediate bootstrapping choice. These findings support a central lesson: increasing n from one can improve learning, but the best choice is not necessarily the largest value tested.
The Step Size α
The learning-rate parameter α affects the rate of learning. It should therefore be considered when comparing different values of n. In the reported Mountain Car comparison, a good step size was α = 0.5/8 for n = 1 and α = 0.3/8 for n = 8. These values belong to that specific experiment. They are not universal settings for every task or every value of n.
| n | Reported α | Context |
|---|---|---|
| n = 1 | 0.5/8 | Good step size in the reported Mountain Car comparison. |
| n = 8 | 0.3/8 | Good step size in the reported Mountain Car comparison. |
Why the Middle Can Win
The Mountain Car results illustrate a trade-off. With a short look ahead, the update can use later estimates quickly, but it incorporates less subsequent experience before the estimate is adjusted. With a longer look ahead, more subsequent experience can contribute, and the reported n = 8 comparison outperformed n = 1. However, the detailed study found n = 4 to be best. This is why an intermediate amount of bootstrapping can be preferable: it can avoid relying too narrowly on a short horizon without assuming that the longest tested horizon is always best.
Treat n and α as related experimental choices. The reported results compare particular settings, so performance claims should always be tied to the task and step size used rather than presented as universal rules.
Check Your Understanding
Explain, in your own words, why n = 8 can outperform n = 1 in the reported Mountain Car comparison while n = 4 can still be identified as the best choice in a more detailed study.
Hints
- Start with what n controls in an update.
- Distinguish the reported comparison from the more detailed study.
- Mention the role of later estimates and the amount of subsequent experience.
A learner says, “Because n = 8 uses more subsequent experience than n = 1, n = 8 must always be the best choice.” Correct the statement using the reported Mountain Car findings and the warning about α.
Hints
- The source reports n = 4 as best in a more detailed study.
- The α values were specific to the reported experiment.
Key Takeaways
- n-step semi-gradient Sarsa combines a semi-gradient Sarsa control method with n-step bootstrapping.
- Bootstrapping lets a later estimate contribute information to the revision of an earlier estimate.
- The value of n controls how far subsequent experience and later estimate information contribute before an update is formed.
- In the reported Mountain Car comparison, n = 8 learned faster and reached better asymptotic performance than n = 1, while a more detailed study found n = 4 to be best.
- α affects the rate of learning, and the reported α values were specific to the experiment rather than universal settings.
Key Takeaways
- n-step semi-gradient Sarsa joins semi-gradient Sarsa control with n-step bootstrapping.
- Bootstrapping transfers information from later estimates backward to revise earlier estimates.
- Larger n allows more subsequent experience to contribute, but the largest n is not always best.
- The reported Mountain Car comparison favored n = 8 over n = 1, while a more detailed study found n = 4 performed best.
- The step size α changes the rate of learning and must be interpreted in the context of the experiment.