Concepts / Tabular TD(0) for Estimating Value Functions

Tabular TD(0) for Estimating Value Functions

The TD error is the difference between the current estimate V(S_t) and the one-step estimate R_t+1 + γV(S_t+1).

  • Programming

A Better Estimate After One Transition

At time t, an agent has a current estimate for the value of its state, written V(S_t). The agent then receives the next reward, R_t+1, and observes the next state, S_t+1. These two new pieces of information let it construct a one-step better estimate. Tabular TD(0) focuses on the difference between the old estimate and this one-step estimate.

The TD error is the difference between the one-step estimate, R_t+1 + γV(S_t+1), and the current estimate, V(S_t). It is commonly written as δ_t = R_t+1 + γV(S_t+1) − V(S_t).

The subscript t identifies the estimate being evaluated: δ_t describes the error in V(S_t). It does not mean that every piece of information needed to calculate δ_t was available at time t.

When the TD Error Becomes Available

The calculation of δ_t needs three ingredients: the current estimate V(S_t), the next reward R_t+1, and the estimate of the next state V(S_t+1). The current estimate is available at time t, but the reward and next state arrive only after the transition. Therefore, the TD error associated with time t becomes available at time t+1.

current estimateobservecombine informationTime tS_t and V(S_t)Transitionnew reward and state arriveTime t+1R_t+1 and S_t+1δ_tnow computable
Why can the TD error for time t only be computed after the transition to S_t+1 and receipt of R_t+1?

Building the One-Step Target

The one-step estimate has two parts. The first is the next reward, R_t+1. The second is the discounted estimate of the next state, γV(S_t+1). Adding them produces the sample backup target R_t+1 + γV(S_t+1). This target is better informed than V(S_t) because it includes information from the transition that has just occurred.

addaddcompare withcompare withR_t+1next rewardR_t+1 + γV(S_t+1)one-step targetV(S_t)current estimateδ_tdifferenceγV(S_t+1)discounted next-stateestimate
How do R_t+1 and the estimated value V(S_t+1) combine to form the one-step target for updating V(S_t)?

Comparing a Current Estimate with a One-Step Target

Suppose V(S_t) is 4, the next reward R_t+1 is 2, γ is 0.5, and V(S_t+1) is 6. Find the one-step target and the TD error.

Form the next-state contribution: The discounted estimate of the next state is γV(S_t+1) = 0.5 × 6 = 3.

Form the one-step target: Add the next reward to the discounted next-state estimate: R_t+1 + γV(S_t+1) = 2 + 3 = 5.

Compare the target with the current estimate: The TD error is the one-step target minus the current estimate: δ_t = 5 − 4 = 1.

The one-step target is 5, and the TD error is 1. The target is one unit higher than the current estimate.

QuantityWhat it usesRole
Current estimate V(S_t)Information represented before the new transitionEstimate being evaluated
One-step estimate R_t+1 + γV(S_t+1)Next reward and discounted next-state estimateBetter estimate based on one transition
TD error δ_tDifference between the one-step estimate and current estimateMeasures the discrepancy between the two estimates

Two Estimates, Two References

The TD error and the Monte Carlo error are not based on the same reference. TD error compares V(S_t) with a one-step estimate made from the next reward and the estimated value of the next state. Monte Carlo error instead uses the actual return. The source gives the important relationship that Monte Carlo error can be written as a sum of TD errors.

comparecomparecomparecompareV(S_t)current estimateδ_tone-step differenceV(S_t)current estimateMonte Carlo errorreturn-based differenceR_t+1 + γV(S_t+1)one-step estimateActual returnfull return
What is the difference between the current estimate V(S_t) and the one-step better estimate R_t+1 + γV(S_t+1)?

Across successive time steps, the TD errors form a sequence. Their sum relates the initial value estimate to the actual return, which is why the Monte Carlo error can be represented as a sum of TD errors. A single TD error looks only one step ahead; the sequence connects those local one-step differences to the full return.

compare at tcontinuecontinuesumV(S_t)initial estimateδ_tfirst one-step differenceδ_t+1next one-step difference…successive TD errorsMonte Carlo errorsum of TD errors
How do TD errors across successive time steps combine to relate the current value estimate to the full Monte Carlo return?

Reading the Value Table

In the tabular setting, each state has a value estimate that can be selected when that state is encountered. For the transition beginning in S_t, the relevant current table entry is V(S_t). The next-state entry V(S_t+1) contributes to the one-step target after it is discounted. The TD error describes the difference between the target and the selected current estimate.

lookuplookupcomparediscounted contributionS_tcurrent stateV(S_t)selected table entryδ_tdifference from one-steptargetS_t+1next stateV(S_t+1)next-state estimate
Which table entry is selected for S_t, and how does the TD error relate the selected entry to the one-step target?

The TD error identifies how the current state's estimate differs from the one-step target. The definition supplied here explains the discrepancy; it does not require replacing the current estimate with the target in one step.

Mistakes with TD Timing

  • Treating δ_t as computable before the transition from S_t.

    The TD error also requires the next reward R_t+1 and the next-state estimate V(S_t+1).

    Fix: Wait until the next reward has been received and S_t+1 has been observed.

  • Using the current estimate as the one-step target.

    The one-step target is formed from R_t+1 and γV(S_t+1).

    Fix: Construct R_t+1 + γV(S_t+1), then compare it with V(S_t).

  • Confusing TD error with Monte Carlo error.

    TD error uses a one-step estimate, whereas Monte Carlo error uses the actual return.

    Fix: Identify the reference first: one-step sample backup for TD, actual return for Monte Carlo.

  • Reading the subscript t as the time when all calculation inputs are available.

    The subscript identifies the estimate V(S_t), but next-step observations are needed for the calculation.

    Fix: Associate δ_t with V(S_t), while remembering that it becomes available at time t+1.

Check Your Understanding

EASY

An agent has a current estimate V(S_t) = 7. It receives R_t+1 = 1, uses γ = 0.5, and has V(S_t+1) = 8. Find the one-step target and the TD error. Then state when this TD error becomes available.

Hints
  • First calculate γV(S_t+1).
  • Add that result to R_t+1 to form the one-step target.
  • Subtract V(S_t) from the target.
  • The required reward and next-state observation arrive at the next time step.

Practice Solution

Use V(S_t) = 7, R_t+1 = 1, γ = 0.5, and V(S_t+1) = 8.

Discount the next-state estimate: γV(S_t+1) = 0.5 × 8 = 4.

Form the one-step target: R_t+1 + γV(S_t+1) = 1 + 4 = 5.

Calculate the TD error: δ_t = 5 − 7 = −2.

Identify availability: The error becomes available at time t+1 because R_t+1 and S_t+1 are needed.

The one-step target is 5, the TD error is −2, and δ_t becomes available at time t+1.

The Essential Distinctions

  1. The TD error compares the current estimate V(S_t) with the one-step estimate R_t+1 + γV(S_t+1).
  2. The one-step estimate combines the next reward with the discounted estimate of the next state.
  3. Although δ_t describes the error in V(S_t), it becomes available at time t+1 because next-step information is required.
  4. Monte Carlo error uses the actual return, while TD error uses a one-step sample backup.
  5. A Monte Carlo error can be written as a sum of TD errors across successive time steps.

Key Takeaways

  • TD(0) evaluates a current value estimate using information from one transition.
  • The one-step target is R_t+1 + γV(S_t+1), and the TD error compares that target with V(S_t).
  • The error labeled δ_t is computed at time t+1 because the next reward and next state are needed.
  • Monte Carlo error uses the actual return and can be expressed as a sum of TD errors.