Concepts / one-step TD backup

one-step TD backup

The λ-return is a normalized weighted average of n-step backups.

  • Programming

From One Backup to Many

A one-step TD backup looks ahead one step. An n-step backup looks ahead a chosen number of steps. The λ-return does not select only one look-ahead length. Instead, it combines the available n-step returns into one return, giving each return a different weight.

The central idea is a weighted average: shorter and longer look-ahead returns all contribute, but their contributions are controlled by λ.

contributescontributescontributesOne-step returnweight 1 − λλ-returnnormalized weighted averageTwo-step returnweight (1 − λ)λLater returnsweights continue fading
How are the one-step, two-step, and later n-step returns combined to produce the λ-return?

The Weighting Pattern

The weights follow a repeated pattern. The one-step return receives weight 1 − λ. To obtain the next weight, multiply the previous weight by λ. Therefore, the two-step return receives (1 − λ)λ, the three-step return receives (1 − λ)λ², and the pattern continues in the same way.

multiply by λmultiply by λcontinue multiplying by λOne-step1 − λTwo-step(1 − λ)λThree-step(1 − λ)λ²Later return(1 − λ)λᵏ
How does each n-step return map to its λ-dependent weight?

Calculating the First Three Weights

Suppose λ = 0.5. Calculate the weights for the one-step, two-step, and three-step returns.

One-step return: Use 1 − λ: 1 − 0.5 = 0.5.

Two-step return: Multiply the first weight by λ: 0.5 × 0.5 = 0.25.

Three-step return: Multiply the second weight by λ again: 0.25 × 0.5 = 0.125.

Interpretation: The one-step return contributes more than the two-step return, and the two-step return contributes more than the three-step return.

The first three weights are 0.5, 0.25, and 0.125.

A Backup Target in Motion

The overall backup changes as λ changes because λ controls how the contribution is distributed across the n-step returns. A small λ places most of the emphasis on the first return. A larger λ allows later returns to receive more of the mixture.

later returns gain weightmixture reaches last returnλ = 0one-step TD onlyIntermediate λone-step and later returnsλ = 1Monte Carlo backup
What changes in the mixture of n-step returns as λ moves from zero toward one?
Value of λBackup behaviorMain contribution
0The overall backup keeps only its first componentOne-step TD backup
Between 0 and 1The backup combines first and later n-step returnsA fading mixture of returns
1The overall backup reduces to its last componentMonte Carlo backup

TD(λ) uses this combined return to update estimates. At λ = 0, there is no contribution from later returns, so the result is the one-step TD backup. At λ = 1, the result reduces to the last return, described in the source as the Monte Carlo backup.

Reading the One-Step Backup

The one-step TD backup is the first component in the λ-return. Conceptually, it uses the current value estimate together with the immediate reward and the value associated with the next state. The λ-return extends this idea by adding longer look-ahead returns instead of relying only on this first component.

part of update contextlooked at immediatelyone-step look-aheadCurrent valueestimateOne-step TD backupImmediate rewardNext-state value
How do the current estimate, immediate reward, and next-state value contribute to the one-step backup?

In the λ-return, this one-step backup is not discarded. It receives the first and largest weight, 1 − λ, before later returns are added with progressively smaller weights.

Terminal States and Later Returns

A terminal state changes how later look-ahead returns behave. Once a terminal state has been reached, all later n-step returns are equal to G_t. Extending the look-ahead beyond the terminal state therefore does not create different return values after that point.

Weights After Termination

An episode reaches a terminal state, and the later n-step returns are all G_t. What does the weighting pattern still do?

Before termination: The one-step, two-step, and other available returns can represent different look-ahead choices, each with its own λ-dependent weight.

At and after termination: Once the terminal state has been reached, all later n-step returns are equal to G_t.

Effect on the combination: The weights may continue as part of the weighting pattern, but the later equal return values do not represent different numerical returns.

Termination makes all later n-step return values equal to G_t, even though the weighting rule continues to describe their contributions.

Common Weighting Mistakes

  • Treating every n-step return as equally important

    The λ-return is a weighted average, and its weights follow the pattern 1 − λ, (1 − λ)λ, (1 − λ)λ², and so on.

    Fix: Start with 1 − λ and multiply by λ for each additional step.

  • Using λ itself as the one-step weight

    The one-step weight is 1 − λ. In this particular numerical example the values happen to match, but that is not the general rule.

    Fix: Always calculate the first weight as 1 − λ before generating later weights.

  • Assuming a larger λ gives more weight to the one-step return

    Each next weight is obtained by multiplying by λ, so increasing λ distributes more of the backup across later returns.

    Fix: Remember that λ controls how much of the overall backup reaches later look-ahead lengths.

  • Ignoring the endpoints

    At λ = 0 the backup keeps only the one-step TD component, while at λ = 1 it reduces to the Monte Carlo backup.

    Fix: Check the endpoint behavior before interpreting an intermediate value of λ.

Practice with Selected Weights

EASY

Let λ = 0.25. Calculate the weights for the one-step, two-step, three-step, and four-step returns.

Hints
  • Calculate the first weight as 1 − λ.
  • Multiply each weight by λ to obtain the next weight.
  • Check that each later weight is smaller than the previous one.

Checking the Practice Calculation

For λ = 0.25, calculate the first four weights.

One-step return: 1 − 0.25 = 0.75.

Two-step return: 0.75 × 0.25 = 0.1875.

Three-step return: 0.1875 × 0.25 = 0.046875.

Four-step return: 0.046875 × 0.25 = 0.01171875.

The first four weights are 0.75, 0.1875, 0.046875, and 0.01171875.

MEDIUM

Compare λ = 0 with λ = 1. State which n-step return controls the overall backup in each case, and name the corresponding backup behavior.

Hints
  • At λ = 0, evaluate the first component of the weighting pattern.
  • At λ = 1, use the endpoint description of the overall backup.
  • Connect the two endpoints to one-step TD and Monte Carlo.

The Backup in One View

  1. The λ-return combines n-step returns into one normalized weighted average. Its weights begin with 1 − λ, then continue as (1 − λ)λ, (1 − λ)λ², and so on. Each additional look-ahead step multiplies the previous weight by λ. When λ = 0, TD(λ) uses the one-step TD backup; when λ = 1, it reduces to the Monte Carlo backup. After a terminal state, all later n-step returns are equal to G_t.

Key Takeaways

  • The λ-return is a normalized weighted average of n-step returns.
  • The weights follow the sequence 1 − λ, (1 − λ)λ, (1 − λ)λ², and continue by multiplying by λ.
  • λ = 0 gives the one-step TD backup, while λ = 1 gives the Monte Carlo backup.
  • A larger λ distributes more of the backup across later n-step returns.
  • After termination, all later n-step returns are equal to G_t.