Concepts / State-Value Prediction

State-Value Prediction

TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.

  • Programming

Two Answers from One Batch

Suppose you receive several complete episodes from an unknown Markov reward process. Your task is to assign a value to each state: the return expected after entering that state. In the source batch, two valid learning methods disagree about state A. Batch Monte Carlo assigns A a value of 0, while batch TD(0) assigns A a value of 3/4. The disagreement is not caused by different episodes. Both methods use the same eight episodes, but they use different evidence from those episodes.

Monte Carlo follows the complete return actually observed after a state. TD(0) uses the reward from the next transition together with the estimated value of the following state.

uses transition evidenceuses complete returnrelies onfollowsSame eight episodesTD(0)A = 3/4Next state estimateB = 3/4Monte CarloA = 0Observed return0 after A
How do TD(0) and Monte Carlo use different evidence from the same episodes to produce different estimates?

The Episode Evidence

The batch contains eight episodes. One episode begins in A, moves to B with reward 0, and then terminates from B with reward 0. The other seven episodes begin in B and terminate immediately. Across all visits to B, six episodes are followed by a return of 1 and two are followed by a return of 0. Therefore, the estimated value of B is 3/4 under both methods.

Reading the batch

Determine the values assigned to B and A by batch TD(0) and batch Monte Carlo.

Find B's observed returns: B is followed by a return of 1 in six episodes and a return of 0 in two episodes. The estimate for B is therefore 3/4.

Trace the visit to A: The only visit to A is followed by a transition to B with reward 0, and the episode then terminates with reward 0.

Apply Monte Carlo reasoning: Monte Carlo uses the complete return actually observed after A. That return is 0, so the batch Monte Carlo estimate for A is 0.

Apply TD(0) reasoning: TD(0) uses the reward of 0 from A to B and the estimated value of the next state B. Since B is estimated as 3/4, batch TD(0) assigns A a value of 3/4.

B is estimated as 3/4 by both methods. A is estimated as 0 by Monte Carlo and 3/4 by TD(0).

average withaverage withaverage withaverage withaverage withaverage withaverage withaverage withB visit 1return 1Bestimated value 3/4B visit 2return 1B visit 3return 1B visit 4return 1B visit 5return 1B visit 6return 1B visit 7return 0B visit 8return 0
What returns are observed after each visit to B, and how do they determine B's estimated value?

Bootstrapping from B

TD(0) does not wait for a complete return before assigning a value to A. It looks at the immediate transition from A to B. The reward on that transition is 0, but the following state B already has an estimated value of 3/4 based on the other episodes. TD(0) treats that next-state estimate as useful evidence about A. In this batch, the resulting value for A is 3/4.

moves withnext statecombines withestimated value informsAcurrent stateReward 0transition evidenceATD(0) estimate 3/4Bestimated value 3/4
How does the reward and the current estimate of the next state combine to update the value of the current state?

Training Fit and Future Prediction

The two estimates answer different practical questions. If the goal is to reproduce the return already recorded after A, Monte Carlo fits that training observation perfectly: the recorded return was 0, and Monte Carlo assigns A a value of 0. TD(0) does not fit that single recorded return as closely, because it assigns A a value of 3/4. However, the goal of state-value prediction is not merely to memorize the returns in one batch. It is to predict what will happen in future episodes.

The transition from A to B provides information about the process. Because the process is Markov, the future depends on the current state rather than on the earlier history once the current state is known. Therefore, after the process reaches B, the earlier fact that the episode came from A does not remove the usefulness of B's observed value. In the given example, this makes the TD(0) estimate of 3/4 for A the estimate expected to have lower error on future data.

leads topredictssupportsEarlier historyhow B was reachedCurrent state Bstate informationFuture returnpredicted from BTD(0) estimateA = 3/4
How does the current state contain enough information to use the next state's value without needing the full history?
fitsinformssupportsRecorded returnafter A0Monte CarloA = 0A to Breward 0Future predictionA = 3/4TD(0)uses B = 3/4
What is the difference between choosing values that fit observed returns and choosing values that predict future episodes?

Mistakes About the Estimates

  • Assuming both methods must produce the same value for every state.

    The methods use different evidence: complete observed returns for Monte Carlo, and the next state's estimated value for TD(0).

    Fix: Compare the evidence used by the method before comparing the resulting estimates.

  • Treating TD(0)'s value for A as the complete return observed after A.

    TD(0) bootstraps from the estimated value of B instead of following only the complete observed return.

    Fix: Separate the observed return after A from the estimate inferred through the transition A to B.

  • Choosing the value that fits the recorded batch and calling it automatically the best future prediction.

    Training-data fit and future-data prediction are different goals.

    Fix: Ask whether the task is reproducing observed returns or generalizing to future episodes.

  • Ignoring the Markov property when evaluating the transition from A to B.

    In the given Markov process, the future depends on the current state rather than the earlier history once the current state is known.

    Fix: Use B's estimated value as relevant evidence for the transition from A to B.

When comparing state-value methods, make an evidence table for each visited state. Record the complete return observed after the visit, the next state reached, and the estimate already available for that next state. This keeps Monte Carlo evidence and TD(0) evidence distinct.

Check Your Reasoning

MEDIUM

Using the source batch, explain in your own words why Monte Carlo assigns A a value of 0 while TD(0) assigns A a value of 3/4. Then state which estimate is expected to have lower error on future data and identify the property of the process that supports that conclusion.

Hints
  • Start with the complete return actually observed after A.
  • Then identify the next state after A and its estimated value.
  • Finally, connect the Markov property to prediction from the current state.

What do you think happens?

Before reading the explanation, predict which method assigns A a value of 0 in the source batch.

  • TD(0)
  • Monte Carlo
  • Both methods
  • Neither method
Reveal answer

Answer: Monte Carlo

Monte Carlo follows the complete return actually observed after A, and that return is 0. TD(0) instead uses the transition from A to B together with B's estimated value of 3/4.

What to Remember

  1. Monte Carlo estimates a state's value from the complete returns actually observed after visits to that state.
  2. TD(0) uses the immediate transition and the estimated value of the following state.
  3. In the source batch, B has value 3/4 under both methods, but A has value 0 under Monte Carlo and 3/4 under TD(0).
  4. A value that fits the recorded training return best is not necessarily the value that predicts future episodes best.
  5. Because the process is Markov, the current state contains the relevant information for using B's value when predicting from the transition A to B.

Key Takeaways

  • Monte Carlo and TD(0) can disagree because they use different evidence from the same episodes.
  • Monte Carlo follows complete observed returns, while TD(0) bootstraps from the estimated value of the next state.
  • For the source batch, both methods estimate B as 3/4; Monte Carlo estimates A as 0, and TD(0) estimates A as 3/4.
  • Monte Carlo fits the recorded return after A, but TD(0) uses the Markov transition structure to support prediction of future data.