Bootstrapping in Temporal-Difference Learning
TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.
One Batch, Two Estimates
Suppose you receive several episodes from an unknown Markov reward process. Your task is to assign a value to each state: how much return should be expected after entering that state? The same episodes can support different estimates for state A. Monte Carlo follows the return actually observed after A. TD(0) instead uses the transition from A to B and then relies on the current estimate for B. This difference is the central idea behind bootstrapping.
The Eight-Episode Batch
The source batch contains eight episodes. One episode begins in A, moves to B with reward 0, and then terminates from B with reward 0. The other seven episodes begin in B and terminate immediately. Across all visits to B, six episodes are followed by a return of 1 and two are followed by a return of 0. Therefore, the batch estimate for B is 3/4.
| Observed episode pattern | Count | Return information |
|---|---|---|
| A moves to B, then terminates | 1 | Reward 0 and return 0 |
| B terminates with return 1 | 6 | Return 1 |
| B terminates with return 0 | 2 | Return 0 |
The observations used by the batch estimates
Tracing the TD(0) Estimate for A
Determine the batch TD(0) and Monte Carlo estimates for state A in the eight-episode batch.
Estimate B: Six returns of 1 and two returns of 0 are observed after visits to B, so the batch estimate for B is 3/4.
Follow the transition from A: The episode containing A moves from A to B with reward 0. TD(0) uses this transition and the estimated value of B.
Apply the TD(0) evidence: Because the reward on the transition is 0 and the following state has estimated value 3/4, batch TD(0) gives A the estimate 3/4.
Follow the observed return: Monte Carlo uses the complete return actually observed after A. That return is 0.
Batch TD(0) gives V(A) = 3/4, while batch Monte Carlo gives V(A) = 0. Both methods give V(B) = 3/4.
Observed Returns in Monte Carlo
Monte Carlo estimation waits for the complete observed return after a state is visited. For A in the source episode, the sequence after the visit is a move to B with reward 0 followed by termination from B with reward 0. The return observed after A is therefore 0, so the batch Monte Carlo estimate for A is 0. Monte Carlo is fitting the outcome recorded after that visit.
Prediction Beyond the Training Batch
The disagreement is not merely a disagreement about arithmetic. It is a disagreement about what evidence should generalize. Monte Carlo gives the value that best reproduces the return already recorded after A: its estimate of 0 has no error on that training observation. TD(0) treats the transition from A to B as informative. Since the process is Markov, the future depends on the current state rather than on earlier history once the current state is known. In this example, that makes the estimated value of B useful evidence about A and makes the TD(0) estimate expected to have lower error on future data.
When comparing learning methods, ask whether the goal is to reproduce the observed training returns or to predict returns in future episodes. The method that fits the recorded data best is not automatically the method expected to predict future data best.
Three Ways to Learn
Temporal-difference learning is a class of methods for solving finite Markov decision problems. Its defining combination is that it requires no model and is fully incremental. Requiring no model means that the method does not depend on a complete and accurate model of the environment. Fully incremental means that learning can make progress step by step rather than waiting for a complete outcome before making progress.
| Method | Model requirement | Episode requirement | Incremental computation | Stated strength or weakness |
|---|---|---|---|---|
| Dynamic programming | Requires a complete and accurate model | Not characterized in the source comparison by waiting for complete observed episodes | Not identified as the defining feature | Mathematically well developed |
| Monte Carlo | Does not require a model | Follows complete observed returns | Not well suited to step-by-step incremental computation | Conceptually simple |
| Temporal-difference learning | Does not require a model | Uses the estimated value of the next state | Fully incremental | More complex to analyze |
Common Reasoning Errors
Assuming that TD(0) and Monte Carlo must produce the same value because they use the same episodes.
The methods use different evidence from the episodes.
Fix:
Identify whether the method follows the complete observed return or uses the estimated value of the following state.Treating the TD(0) estimate for A as the directly observed return after A.
TD(0) uses the transition from A to B and the estimate for B.
Fix:
Trace the next state before deciding what evidence TD(0) uses.Calling Monte Carlo's estimate the best possible prediction simply because it fits the recorded episode.
Fitting training data and predicting future data are different goals.
Fix:
State clearly whether you are discussing training-data fit or future prediction.Saying that no model means no information about transitions is used.
No model requirement concerns what must be supplied about the environment, not whether observed experience is ignored.
Fix:
Distinguish a supplied environment model from information obtained through observed episodes.Claiming that temporal-difference learning is always superior.
The source identifies its flexibility and incremental computation, but also says it is more complex to analyze and does not give a universal speed or efficiency ranking.
Fix:
Compare methods by their requirements and trade-offs rather than declaring one universally best.
Check Your Understanding
A learner says: The return after A was 0, so TD(0) must assign A the value 0. Explain why this conclusion is incorrect in the source batch, and identify the value assigned by batch TD(0).
Hints
- First identify the following state after A.
- Then identify the batch estimate for that following state.
- Finally distinguish the complete observed return from the next-state estimate.
The Bootstrapping Pattern
Bootstrapping means that TD(0) makes an estimate for the current state using an estimate for a following state. In the source batch, this gives A the value 3/4 because A leads to B and the batch estimate for B is 3/4. Monte Carlo instead uses the complete observed return after A, which is 0. The two estimates differ because they answer different generalization questions: one follows the recorded outcome, while the other uses transition information that is useful for predicting future episodes.
- Monte Carlo follows complete observed returns, while TD(0) uses the estimated value of the next state.
- In the source batch, both methods give V(B) = 3/4, but TD(0) gives V(A) = 3/4 and Monte Carlo gives V(A) = 0.
- The estimate that fits recorded training data best is not necessarily the estimate that predicts future episodes best.
- The Markov property makes the transition from A to B useful evidence about future returns in the given example.
- Temporal-difference learning requires no model, is fully incremental, and is more complex to analyze than the alternatives described.
Key Takeaways
- TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.
- Monte Carlo uses the complete return observed after a state, whereas TD(0) bootstraps from the estimated value of the following state.
- In the source batch, TD(0) gives A a value of 3/4 and Monte Carlo gives A a value of 0.
- The Markov property supports using the current state and its transition information to predict future episodes.
- Temporal-difference learning combines no model requirement with fully incremental computation, although it is more complex to analyze.