Concepts / n-Step Bootstrapping

n-Step Bootstrapping

The differential n-step return is G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)).

  • Programming

From Current State to Later State

n-step bootstrapping uses information from a later point in the process to construct a return for the state at time t. When function approximation is used, the return is written in differential form. Instead of treating the target as an isolated absolute quantity, the expression combines an estimate of average reward with the change between two estimated state values.

n-step constructionlater-state estimateS(t)V̄(S(t))R̄average-reward estimateS(t+n)V̄(S(t+n))
How do the starting state, the average-reward estimate, and the state n steps later connect across the n-step window?

The role of n is to determine which later state appears in the expression. For example, when n = 3, the later state is S(t+3), and its estimate is V̄(S(t+3)).

The Differential Return

G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t))

Read the expression from left to right. Begin with R̄, an estimate of η(π), the average reward associated with the policy. Then add the estimated value of the state n steps later, V̄(S(t+n)). Finally, subtract the estimated value of the starting state, V̄(S(t)). The result is the differential n-step return G(n)t.

addaddsubtractG(n)tdifferential n-step returnR̄estimate of η(π)V̄(S(t+n))later-state value estimateV̄(S(t))starting-state valueestimate
What does each term in G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)) represent, and how do the terms combine?

Substituting a Three-Step Later State

Write the differential return when n = 3.

Identify the later state: Replacing n with 3 makes the later state S(t+3).

Preserve the structure: The average-reward estimate is added, the later-state estimate is added, and the starting-state estimate is subtracted.

Write the expression: The return becomes G(3)t = R̄ + V̄(S(t+3)) − V̄(S(t)).

G(3)t = R̄ + V̄(S(t+3)) − V̄(S(t))

Why Values Appear as a Difference

With function approximation, the return is expressed in differential form. The value contribution is therefore a change from the starting state's estimate to the later state's estimate: V̄(S(t+n)) − V̄(S(t)). This is the form specified for the n-step return in the source material; it combines the average-reward estimate with the difference between two estimated state values.

subtractaddAbsolute returnnot the displayeddifferential formV̄(S(t))starting estimateDifferential returnR̄ + V̄(S(t+n)) − V̄(S(t))V̄(S(t+n))later estimate
Why does the function-approximation form use a change between two estimated values?

When the Expression Applies

The displayed differential n-step return applies when n is at least 1 and t+n is less than T. The second condition means that the later state used in the expression occurs before time T. Under these conditions, the expression uses S(t+n) and its estimated value.

differs by conditiondiffers by conditionDisplayed formn ≥ 1 and t+n < Tn = 0outside stated conditiont+n ≥ Toutside stated condition
What condition makes the displayed differential n-step return applicable?

From Return to TD Error

The n-step return and the n-step TD error are related but distinct quantities. The return is formed first using G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)). The n-step TD error is then defined using that differential return and the current value estimate. In other words, G(n)t supplies the target used when determining the error; it is not itself the error.

comparecompareG(n)tdifferential n-step returnn-step TD errordefined using the returnV̄(S(t))current value estimate
How is the n-step TD error obtained after the differential n-step return has been formed?

Use this order: first form the differential n-step return, then use that return to define the n-step TD error. Mixing the two names hides the learning sequence.

Common Mistakes

  • Treating G(n)t as the n-step TD error.

    The return and the TD error are related but are not the same named quantity.

    Fix: Form G(n)t first, then define the n-step TD error using that return.

  • Using the wrong later state.

    The later state must be S(t+n), so changing n changes which later estimate appears.

    Fix: Substitute the selected n carefully before writing the expression.

  • Dropping the starting-state term.

    The differential form includes both estimated state values, with the starting estimate subtracted.

    Fix: Keep the complete structure R̄ + V̄(S(t+n)) − V̄(S(t)).

  • Ignoring the time condition.

    Those are the stated conditions for the displayed formula.

    Fix: Check both inequalities before using the expression.

Practice Check

EASY

Suppose the step count is n = 4. Write the differential n-step return, identify the starting-state estimate, identify the later-state estimate, and state the condition that must hold for the displayed form to apply.

Hints
  • Replace n with 4 in S(t+n).
  • The starting-state estimate is the term that is subtracted.
  • State both parts of the condition: one involving n and one involving t+n.

Practice Check Solution

Write the differential return for n = 4 and identify its components.

Substitute the step count: The later state becomes S(t+4).

Keep the differential structure: Add R̄ and V̄(S(t+4)), then subtract V̄(S(t)).

Check applicability: The displayed form requires n ≥ 1 and t+n < T. With n = 4, this requires 4 ≥ 1 and t+4 < T.

G(4)t = R̄ + V̄(S(t+4)) − V̄(S(t)); the starting-state estimate is V̄(S(t)), and the later-state estimate is V̄(S(t+4)).

Key Takeaways

  1. With function approximation, the n-step return is written in differential form.
  2. The differential return is G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)).
  3. R̄ estimates η(π); V̄(S(t+n)) is the later-state estimate; and V̄(S(t)) is the starting-state estimate that is subtracted.
  4. The displayed expression applies when n ≥ 1 and t+n < T.
  5. The n-step TD error is defined using the differential n-step return, so form the return before forming the error.

Key Takeaways

  • The differential n-step return combines an average-reward estimate with a difference between later and starting state-value estimates.
  • The step count n determines which later state appears in the return.
  • The displayed formula requires n ≥ 1 and t+n < T.
  • The n-step TD error is defined after the differential n-step return has been constructed.