n-Step TD Prediction Methods
The full return is the complete discounted reward sequence through episode termination.
From Complete Return to Shorter Target
An episode can produce a sequence of rewards after a particular time step. The full return uses the complete discounted reward sequence from that time step through episode termination. An n-step return uses only the rewards observed during the next n steps and then estimates the remaining part. The central idea is to replace a long, completely observed target with a shorter target that combines observed rewards and a value estimate.
An n-step return is not a completely different target from the full return. It is an approximation to the full return: the first n rewards are explicit, while the later terms are represented by a value estimate.
Following the n-Step Window
Start at time t and look forward one step at a time. The n-step window includes the reward from the first step after t, the reward from the second step after t, and so on until n rewards have been observed. Those rewards form the explicit part of the n-step return. Rewards after that window are not individually included in the target. Instead, the value estimate for the state reached after n steps represents the missing later terms.
Selecting the Explicit Rewards
Suppose the next four rewards after time t are observed, but the method uses a 2-step return.
Locate the starting time: Begin with the state at time t.
Count two steps: The rewards from the first and second steps after t are included explicitly.
Stop the explicit sequence: The third and fourth rewards are beyond the 2-step window, so they are not written as individually observed terms in the n-step target.
Represent what remains: A value estimate at the state reached after the second step represents the missing later terms.
The 2-step return contains two explicit rewards followed by a discounted value estimate for the remaining part.
The Bootstrap Correction
Stopping after n rewards would leave out the rewards that occur later in the episode. The n-step method addresses this by looking at the state reached after n steps and using its value estimate as a stand-in for those missing later terms. Because this estimate begins n steps after the starting time, its contribution is discounted by γ^n. Thus, the n-step target has two parts: the explicitly observed discounted rewards and the discounted value estimate for the remaining sequence.
Separating Observed and Estimated Parts
Consider a 3-step return beginning at time t.
Include three rewards: The rewards observed on steps one, two, and three after time t belong explicitly to the n-step return.
Find the boundary state: After the third step, identify the state reached at the end of the 3-step window.
Estimate the remainder: Use the value estimate of that boundary state to represent the later discounted rewards.
Discount the estimate: The value estimate is multiplied by γ^3 because it stands three steps away from the starting time.
A 3-step return is the discounted contribution of the first three rewards plus γ^3 times the value estimate at the state reached after three steps.
Full Return and n-Step Approximation
The full return continues explicitly through episode termination, so it contains the entire discounted reward sequence from time t onward. The n-step return stops writing rewards after n steps. It then uses the value estimate at the reached state to approximate what the full return would have received later. The two targets therefore share the same beginning, but the n-step return replaces the later explicit sequence with an estimate.
| Feature | Full return | n-step return |
|---|---|---|
| Reward sequence | Uses the complete sequence through episode termination | Uses rewards observed during the next n steps |
| Later terms | Remain explicit in the return | Are represented by a value estimate |
| Role of γ^n | Not the defining correction described for a truncated window | Discounts the value estimate at the n-step boundary |
| Relationship | Complete target | Approximation to the full return |
Termination and Exact Equality
The n-step return equals the ordinary full return when the n-step window reaches episode termination. In that case, there are no later rewards left to represent with a value estimate. The truncated sequence is not actually missing a later part, so the n-step target contains the same complete reward sequence as the full return.
An n-Step Window at the Episode End
Suppose an episode ends exactly after the nth step of the window.
Observe the n rewards: The n-step method observes every reward from the starting time through the end of the window.
Check for later terms: Because the episode terminates at that boundary, there are no later rewards beyond the observed sequence.
Remove the missing-part correction: There is no remaining sequence for a value estimate to represent.
The n-step return is equal to the ordinary full return because the n-step window reaches episode termination.
Common Reasoning Errors
Treating the n-step return as if it includes every reward through episode termination.
The n-step return truncates the explicit reward sequence after n steps.
Fix:
Include only the next n observed rewards explicitly, then use the value estimate for the later terms.Using the value estimate to replace the first n rewards.
The estimate represents the missing later part, not the already observed rewards.
Fix:
Keep the first n discounted rewards and add the discounted value estimate for the remainder.Forgetting the γ^n discount on the value estimate.
The source definition states that the correction is discounted by γ^n.
Fix:
Discount the value estimate by γ^n before combining it with the explicit reward sequence.Assuming an n-step return is always different from the full return.
When the window reaches termination, no later terms are missing.
Fix:
Recognize that the n-step and full returns are equal when the n-step window reaches episode termination.
When analyzing an n-step target, mark the boundary first. Count exactly n steps from time t, list the rewards up to that boundary, and then ask whether the episode has already terminated. If it has not, the value estimate at the reached state represents the later terms and must receive the γ^n discount.
Practice Check
A learner is forming a 4-step return from time t. The episode continues beyond the fourth step. Identify which rewards are included explicitly, what represents the later terms, and what discount is applied to that estimate. Then state whether this target is the full return or an approximation to it.
Hints
- Count four rewards beginning with the reward from the first step after time t.
- Locate the state reached after the fourth step.
- The missing later terms are represented by that state's value estimate.
- The correction is discounted by γ raised to the number of steps in the window.
What do you think happens?
If an episode terminates exactly at the end of a 4-step window, does the 4-step return still need a value estimate for later rewards?
Reveal answer
Answer: No, because there are no later terms after termination.
When the n-step window reaches episode termination, the observed sequence already extends through the end of the episode. Therefore, the n-step return equals the ordinary full return and has no missing later part to estimate.
Essential Distinctions
- The full return is the complete discounted reward sequence from time t through episode termination.
- The n-step return includes the rewards observed during the next n steps explicitly.
- The value estimate at the state reached after n steps represents the missing later terms.
- That estimate is discounted by γ^n because it begins n steps after the starting time.
- The n-step return equals the full return when its window reaches episode termination.
Key Takeaways
- The full return follows discounted rewards all the way to episode termination.
- An n-step return follows only the next n rewards explicitly.
- A value estimate at the state reached after n steps stands in for the later missing terms.
- The bootstrap correction is discounted by γ^n.
- If the n-step window reaches termination, the n-step return and full return are equal.