one-step TD backup
The λ-return is a normalized weighted average of n-step backups.
From One Backup to Many
A one-step TD backup looks ahead one step. An n-step backup looks ahead a chosen number of steps. The λ-return does not select only one look-ahead length. Instead, it combines the available n-step returns into one return, giving each return a different weight.
The central idea is a weighted average: shorter and longer look-ahead returns all contribute, but their contributions are controlled by λ.
The Weighting Pattern
The weights follow a repeated pattern. The one-step return receives weight 1 − λ. To obtain the next weight, multiply the previous weight by λ. Therefore, the two-step return receives (1 − λ)λ, the three-step return receives (1 − λ)λ², and the pattern continues in the same way.
Calculating the First Three Weights
Suppose λ = 0.5. Calculate the weights for the one-step, two-step, and three-step returns.
One-step return: Use 1 − λ: 1 − 0.5 = 0.5.
Two-step return: Multiply the first weight by λ: 0.5 × 0.5 = 0.25.
Three-step return: Multiply the second weight by λ again: 0.25 × 0.5 = 0.125.
Interpretation: The one-step return contributes more than the two-step return, and the two-step return contributes more than the three-step return.
The first three weights are 0.5, 0.25, and 0.125.
A Backup Target in Motion
The overall backup changes as λ changes because λ controls how the contribution is distributed across the n-step returns. A small λ places most of the emphasis on the first return. A larger λ allows later returns to receive more of the mixture.
| Value of λ | Backup behavior | Main contribution |
|---|---|---|
| 0 | The overall backup keeps only its first component | One-step TD backup |
| Between 0 and 1 | The backup combines first and later n-step returns | A fading mixture of returns |
| 1 | The overall backup reduces to its last component | Monte Carlo backup |
TD(λ) uses this combined return to update estimates. At λ = 0, there is no contribution from later returns, so the result is the one-step TD backup. At λ = 1, the result reduces to the last return, described in the source as the Monte Carlo backup.
Reading the One-Step Backup
The one-step TD backup is the first component in the λ-return. Conceptually, it uses the current value estimate together with the immediate reward and the value associated with the next state. The λ-return extends this idea by adding longer look-ahead returns instead of relying only on this first component.
In the λ-return, this one-step backup is not discarded. It receives the first and largest weight, 1 − λ, before later returns are added with progressively smaller weights.
Terminal States and Later Returns
A terminal state changes how later look-ahead returns behave. Once a terminal state has been reached, all later n-step returns are equal to G_t. Extending the look-ahead beyond the terminal state therefore does not create different return values after that point.
Weights After Termination
An episode reaches a terminal state, and the later n-step returns are all G_t. What does the weighting pattern still do?
Before termination: The one-step, two-step, and other available returns can represent different look-ahead choices, each with its own λ-dependent weight.
At and after termination: Once the terminal state has been reached, all later n-step returns are equal to G_t.
Effect on the combination: The weights may continue as part of the weighting pattern, but the later equal return values do not represent different numerical returns.
Termination makes all later n-step return values equal to G_t, even though the weighting rule continues to describe their contributions.
Common Weighting Mistakes
Treating every n-step return as equally important
The λ-return is a weighted average, and its weights follow the pattern 1 − λ, (1 − λ)λ, (1 − λ)λ², and so on.
Fix:
Start with 1 − λ and multiply by λ for each additional step.Using λ itself as the one-step weight
The one-step weight is 1 − λ. In this particular numerical example the values happen to match, but that is not the general rule.
Fix:
Always calculate the first weight as 1 − λ before generating later weights.Assuming a larger λ gives more weight to the one-step return
Each next weight is obtained by multiplying by λ, so increasing λ distributes more of the backup across later returns.
Fix:
Remember that λ controls how much of the overall backup reaches later look-ahead lengths.Ignoring the endpoints
At λ = 0 the backup keeps only the one-step TD component, while at λ = 1 it reduces to the Monte Carlo backup.
Fix:
Check the endpoint behavior before interpreting an intermediate value of λ.
Practice with Selected Weights
Let λ = 0.25. Calculate the weights for the one-step, two-step, three-step, and four-step returns.
Hints
- Calculate the first weight as 1 − λ.
- Multiply each weight by λ to obtain the next weight.
- Check that each later weight is smaller than the previous one.
Checking the Practice Calculation
For λ = 0.25, calculate the first four weights.
One-step return: 1 − 0.25 = 0.75.
Two-step return: 0.75 × 0.25 = 0.1875.
Three-step return: 0.1875 × 0.25 = 0.046875.
Four-step return: 0.046875 × 0.25 = 0.01171875.
The first four weights are 0.75, 0.1875, 0.046875, and 0.01171875.
Compare λ = 0 with λ = 1. State which n-step return controls the overall backup in each case, and name the corresponding backup behavior.
Hints
- At λ = 0, evaluate the first component of the weighting pattern.
- At λ = 1, use the endpoint description of the overall backup.
- Connect the two endpoints to one-step TD and Monte Carlo.
The Backup in One View
- The λ-return combines n-step returns into one normalized weighted average. Its weights begin with 1 − λ, then continue as (1 − λ)λ, (1 − λ)λ², and so on. Each additional look-ahead step multiplies the previous weight by λ. When λ = 0, TD(λ) uses the one-step TD backup; when λ = 1, it reduces to the Monte Carlo backup. After a terminal state, all later n-step returns are equal to G_t.
Key Takeaways
- The λ-return is a normalized weighted average of n-step returns.
- The weights follow the sequence 1 − λ, (1 − λ)λ, (1 − λ)λ², and continue by multiplying by λ.
- λ = 0 gives the one-step TD backup, while λ = 1 gives the Monte Carlo backup.
- A larger λ distributes more of the backup across later n-step returns.
- After termination, all later n-step returns are equal to G_t.