Concepts / Bootstrapping in Temporal-Difference Learning

Bootstrapping in Temporal-Difference Learning

TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.

  • Programming

One Batch, Two Estimates

Suppose you receive several episodes from an unknown Markov reward process. Your task is to assign a value to each state: how much return should be expected after entering that state? The same episodes can support different estimates for state A. Monte Carlo follows the return actually observed after A. TD(0) instead uses the transition from A to B and then relies on the current estimate for B. This difference is the central idea behind bootstrapping.

TD evidenceuse next statesupportsobserved after visitsupportsAcurrent stateA to Breward 0V(B) = 3/4estimated next-state valueReturn after A = 0observed outcomeTD(0): 3/4estimate for AMonte Carlo: 0estimate for A
How can TD(0) and Monte Carlo assign different values to the same state when they learn from the same episodes?

The Eight-Episode Batch

The source batch contains eight episodes. One episode begins in A, moves to B with reward 0, and then terminates from B with reward 0. The other seven episodes begin in B and terminate immediately. Across all visits to B, six episodes are followed by a return of 1 and two are followed by a return of 0. Therefore, the batch estimate for B is 3/4.

Observed episode patternCountReturn information
A moves to B, then terminates1Reward 0 and return 0
B terminates with return 16Return 1
B terminates with return 02Return 0

The observations used by the batch estimates

transitionfollowed byestimated asbootstrapsAcurrent state0observed rewardBfollowing state3/4estimate for B3/4batch TD(0) estimate for A
How does the estimated value of the next state flow into the update for the current state?

Tracing the TD(0) Estimate for A

Determine the batch TD(0) and Monte Carlo estimates for state A in the eight-episode batch.

Estimate B: Six returns of 1 and two returns of 0 are observed after visits to B, so the batch estimate for B is 3/4.

Follow the transition from A: The episode containing A moves from A to B with reward 0. TD(0) uses this transition and the estimated value of B.

Apply the TD(0) evidence: Because the reward on the transition is 0 and the following state has estimated value 3/4, batch TD(0) gives A the estimate 3/4.

Follow the observed return: Monte Carlo uses the complete return actually observed after A. That return is 0.

Batch TD(0) gives V(A) = 3/4, while batch Monte Carlo gives V(A) = 0. Both methods give V(B) = 3/4.

Observed Returns in Monte Carlo

Monte Carlo estimation waits for the complete observed return after a state is visited. For A in the source episode, the sequence after the visit is a move to B with reward 0 followed by termination from B with reward 0. The return observed after A is therefore 0, so the batch Monte Carlo estimate for A is 0. Monte Carlo is fitting the outcome recorded after that visit.

moves toreceivesthenobserved returnAstate visitBnext state0rewardTerminationreturn 0V(A) = 0Monte Carlo estimate
After a state is visited, how does the sequence of rewards that actually follows it become an estimate of that state's value?

Prediction Beyond the Training Batch

The disagreement is not merely a disagreement about arithmetic. It is a disagreement about what evidence should generalize. Monte Carlo gives the value that best reproduces the return already recorded after A: its estimate of 0 has no error on that training observation. TD(0) treats the transition from A to B as informative. Since the process is Markov, the future depends on the current state rather than on earlier history once the current state is known. In this example, that makes the estimated value of B useful evidence about A and makes the TD(0) estimate expected to have lower error on future data.

fitsinterpreted throughsupportspredictsObserved returnafter A0Monte CarloV(A) = 0A to Btransition evidenceMarkov propertycurrent state isinformativeTD(0)V(A) = 3/4Future episodespredicted return
What is the difference between choosing values that fit the returns already observed and choosing values that predict what will happen in future episodes?

When comparing learning methods, ask whether the goal is to reproduce the observed training returns or to predict returns in future episodes. The method that fits the recorded data best is not automatically the method expected to predict future data best.

Three Ways to Learn

Temporal-difference learning is a class of methods for solving finite Markov decision problems. Its defining combination is that it requires no model and is fully incremental. Requiring no model means that the method does not depend on a complete and accurate model of the environment. Fully incremental means that learning can make progress step by step rather than waiting for a complete outcome before making progress.

requiresusesusessupportshasDynamic programmingmodel requiredComplete accuratemodelrequired by DPStep-by-step progressfully incrementalMonte Carlono modelComplete episodeused by Monte CarloAnalysis complexityhigher for TDTemporal differenceno modelEstimated next stateused by TD
How do dynamic programming, Monte Carlo, and temporal-difference learning differ in their use of a model, complete episodes, bootstrapping, and incremental updates?
MethodModel requirementEpisode requirementIncremental computationStated strength or weakness
Dynamic programmingRequires a complete and accurate modelNot characterized in the source comparison by waiting for complete observed episodesNot identified as the defining featureMathematically well developed
Monte CarloDoes not require a modelFollows complete observed returnsNot well suited to step-by-step incremental computationConceptually simple
Temporal-difference learningDoes not require a modelUses the estimated value of the next stateFully incrementalMore complex to analyze

Common Reasoning Errors

  • Assuming that TD(0) and Monte Carlo must produce the same value because they use the same episodes.

    The methods use different evidence from the episodes.

    Fix: Identify whether the method follows the complete observed return or uses the estimated value of the following state.

  • Treating the TD(0) estimate for A as the directly observed return after A.

    TD(0) uses the transition from A to B and the estimate for B.

    Fix: Trace the next state before deciding what evidence TD(0) uses.

  • Calling Monte Carlo's estimate the best possible prediction simply because it fits the recorded episode.

    Fitting training data and predicting future data are different goals.

    Fix: State clearly whether you are discussing training-data fit or future prediction.

  • Saying that no model means no information about transitions is used.

    No model requirement concerns what must be supplied about the environment, not whether observed experience is ignored.

    Fix: Distinguish a supplied environment model from information obtained through observed episodes.

  • Claiming that temporal-difference learning is always superior.

    The source identifies its flexibility and incremental computation, but also says it is more complex to analyze and does not give a universal speed or efficiency ranking.

    Fix: Compare methods by their requirements and trade-offs rather than declaring one universally best.

Check Your Understanding

MEDIUM

A learner says: The return after A was 0, so TD(0) must assign A the value 0. Explain why this conclusion is incorrect in the source batch, and identify the value assigned by batch TD(0).

Hints
  • First identify the following state after A.
  • Then identify the batch estimate for that following state.
  • Finally distinguish the complete observed return from the next-state estimate.
identifiesbecomes unnecessarysupportscompared withCurrent stateAEarlier historynot needed once state isknownMarkov propertyfuture depends on currentstateV(B) = 3/4TD(0) evidenceObserved return 0one training outcome
Why does knowing the current state make the following-state estimate more reliable than averaging the limited returns observed in the training episodes?

The Bootstrapping Pattern

Bootstrapping means that TD(0) makes an estimate for the current state using an estimate for a following state. In the source batch, this gives A the value 3/4 because A leads to B and the batch estimate for B is 3/4. Monte Carlo instead uses the complete observed return after A, which is 0. The two estimates differ because they answer different generalization questions: one follows the recorded outcome, while the other uses transition information that is useful for predicting future episodes.

  1. Monte Carlo follows complete observed returns, while TD(0) uses the estimated value of the next state.
  2. In the source batch, both methods give V(B) = 3/4, but TD(0) gives V(A) = 3/4 and Monte Carlo gives V(A) = 0.
  3. The estimate that fits recorded training data best is not necessarily the estimate that predicts future episodes best.
  4. The Markov property makes the transition from A to B useful evidence about future returns in the given example.
  5. Temporal-difference learning requires no model, is fully incremental, and is more complex to analyze than the alternatives described.

Key Takeaways

  • TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.
  • Monte Carlo uses the complete return observed after a state, whereas TD(0) bootstraps from the estimated value of the following state.
  • In the source batch, TD(0) gives A a value of 3/4 and Monte Carlo gives A a value of 0.
  • The Markov property supports using the current state and its transition information to predict future episodes.
  • Temporal-difference learning combines no model requirement with fully incremental computation, although it is more complex to analyze.