Concepts / n-Step TD Error

n-Step TD Error

The differential n-step return is G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)).

  • Programming

From Return to Learning Signal

In n-step bootstrapping, the n-step return supplies the target information needed to form the n-step TD error. When function approximation is used, the return is written in differential form: it combines an estimate of the average reward with the difference between a later estimated state value and the estimated value of the starting state.

The n-step return and the n-step TD error are related, but they are not the same named quantity. First form the differential n-step return. Then define the n-step TD error using that return.

The Differential Return

G⁽ⁿ⁾ₜ = R̄ + V̄(Sₜ₊ₙ) − V̄(Sₜ)

Read the expression from left to right. R̄ is an estimate of η(π), the average reward associated with the policy. V̄(Sₜ₊ₙ) is the estimated value of the later state reached n steps after time t. V̄(Sₜ) is the estimated value of the state at time t. The later-state estimate and the starting-state estimate appear as a difference, so the return is expressed relative to the current estimate.

G⁽ⁿ⁾ₜdifferential return=is formed asR̄average reward estimate+addsV̄(Sₜ₊ₙ)later-state value estimate−subtractsV̄(Sₜ)starting-state valueestimate
How do the average reward, later-state estimate, and starting-state estimate combine in the differential n-step return?

Changing the Step Horizon

Replacing n with 3

Trace the differential return when n = 3, without assigning numerical values.

Identify the later state: With n = 3, the later state is the state at time t+3, so the later estimated value is V̄(Sₜ₊₃).

Keep the average-reward term: The expression still includes R̄, the estimate of η(π).

Keep the starting-state term: The expression still subtracts V̄(Sₜ), the estimated value of the state at time t.

G⁽³⁾ₜ = R̄ + V̄(Sₜ₊₃) − V̄(Sₜ)

Changing n changes which later state appears in the expression. The structure does not otherwise change: the differential return still combines the average-reward estimate, the later-state value estimate, and the starting-state value estimate, provided the calculation ends before time T.

From Return to TD Error

The n-step TD error is defined using the differential n-step return G⁽ⁿ⁾ₜ. The return is formed first, and the resulting quantity supplies the learning signal for the next stage of the algorithm.

definesG⁽ⁿ⁾ₜdifferential returnn-step TD errorlearning signal
What changes when the differential n-step return becomes the n-step TD error?

The Two-Stage Algorithm

The n-step TD error has a procedural role: it prepares the learning signal for the following semi-gradient Sarsa update. The algorithm treats these as two connected but distinct stages. First, calculate the n-step TD error. Next, apply the usual semi-gradient Sarsa update using the result of that calculation.

supplies target informationprovides learning signalchangesDifferential returnG⁽ⁿ⁾ₜn-step TD errorcalculated learning signalSemi-gradient Sarsaupdatenext operationParametersupdated learning system
What happens after the n-step TD error is calculated, and how does it flow into the parameter update?

Keep error calculation and parameter updating conceptually separate when tracing the algorithm. The update must follow the error calculation because the calculated error is the input that makes the next update stage ready to begin.

Mistakes in the Calculation Order

  • Treating the differential n-step return and the n-step TD error as interchangeable names.

    The source distinguishes the return from the error. The return is formed first, and the n-step TD error is then defined using that return.

    Fix: Describe the return as the quantity that supplies the basis or target information for defining the n-step TD error.

  • Applying the semi-gradient Sarsa update before calculating the n-step TD error.

    The source specifies that the n-step TD error is calculated first and that the semi-gradient Sarsa update follows it.

    Fix: Use the order: form the differential return, calculate the n-step TD error, then apply the semi-gradient Sarsa update.

  • Forgetting which state changes when n changes.

    The later-state estimate depends on n. For n = 3, it is V̄(Sₜ₊₃); another n identifies another later state.

    Fix: Replace the later-state index with t+n while keeping the overall differential structure.

Practice Trace

EASY

Trace the high-level operation order for an n-step bootstrapping step. Begin with the differential return, identify what is calculated from it, and name the update that follows.

Hints
  • The differential return contains R̄, V̄(Sₜ₊ₙ), and V̄(Sₜ).
  • The n-step TD error is the stage between the return and the parameter update.
  • The semi-gradient Sarsa update follows the error calculation.
  1. A correct trace is: form the differential n-step return, use it to define the n-step TD error, and then apply the semi-gradient Sarsa update.

Key Takeaways

  • For function approximation, the n-step return is expressed in differential form as G⁽ⁿ⁾ₜ = R̄ + V̄(Sₜ₊ₙ) − V̄(Sₜ).
  • R̄ estimates the average reward, while the two V̄ terms estimate the later and starting states.
  • The displayed differential return applies when n is at least 1 and t+n is less than T.
  • The n-step TD error is defined using the differential n-step return.
  • The algorithm calculates the n-step TD error before applying the semi-gradient Sarsa update.