Concepts / Returns and Monte Carlo Prediction

Returns and Monte Carlo Prediction

TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.

  • Programming

Introduction to Returns and Monte Carlo Prediction

TD(0) and Monte Carlo methods are used to estimate the value of states in a Markov reward process. Although they use the same episodes, they can produce different estimates for the same state. This difference arises because they utilize different evidence from the episodes.

How TD(0) and Monte Carlo Assign Different Values

The key difference between TD(0) and Monte Carlo lies in how they estimate the value of a state. Monte Carlo follows the complete observed returns after visiting a state, whereas TD(0) uses the estimated value of the next state it transitions to.

usesusesobserved returnbootstraps fromMonte CarloTD(0)State A (MC)State A (TD(0))Observed ReturnNext State's Value
Comparing how TD(0) and Monte Carlo estimate state values from the same episodes.

Batch TD(0) Bootstrapping

Batch TD(0) updates the estimate of a state's value by using the estimated value of the next state it transitions to. This process is called bootstrapping.

transitions toused to updateCurrent StateNext StateCurrent State's ValueNext State's Value
How the estimated value of the next state flows into the update for the current state in batch TD(0).

Monte Carlo Estimation

Monte Carlo estimation involves averaging the returns observed after visiting a state. The value of a state is the average return following visits to that state.

followed byfollowed byaveragedaveragedState VisitReturn 1Return 2Average Return
How Monte Carlo estimates a state's value from observed returns after visiting it.

Fitting Training Data vs Predicting Future Data

Monte Carlo is good at fitting the existing training data, as it directly averages the observed returns. However, TD(0) is better at predicting future data because it uses the Markov property to generalize from the transition patterns observed.

fitspredictsMonte CarloTD(0)Training DataFuture Data
The difference between fitting observed returns and predicting future returns.

The Markov Property and TD(0)

The Markov property states that the future depends only on the current state, not on the history leading to it. TD(0) leverages this property by using the next state's value to update the current state's value, making it effective for prediction in Markov processes.

determinesdeterminesCurrent StateNext StateFuture
How the Markov property supports TD(0) estimation.

Common Mistakes

  • Assuming TD(0) and Monte Carlo always give the same estimates.

    TD(0) uses bootstrapping from the next state, while Monte Carlo averages observed returns.

    Fix: Understand that different methods use different evidence from the same episodes.

  • Thinking Monte Carlo is always better because it directly uses observed returns.

    While Monte Carlo fits the training data well, it may not predict future data as effectively as TD(0) in Markov processes.

    Fix: Consider the Markov property and the need for generalization when choosing between methods.

Practice

MEDIUM

Consider a simple Markov reward process with two states, A and B, where A transitions to B with a reward of 0, and B terminates with a reward of 1 or 0. Use both TD(0) and Monte Carlo methods to estimate the value of state A after observing several episodes.

Hints
  • Start by calculating the observed returns for state A.
  • Use the estimated value of B to update the value of A in TD(0).

Summary

  1. TD(0) and Monte Carlo can assign different values to the same state due to their different approaches to estimation.
  2. TD(0) bootstraps from the estimated value of the next state.
  3. Monte Carlo estimates a state's value by averaging the observed returns after visiting it.
  4. The Markov property favors TD(0) for prediction in Markov processes.

Key Takeaways

  • TD(0) and Monte Carlo methods can produce different estimates for the same state due to their different uses of episode data.
  • TD(0) uses bootstrapping from the next state's estimated value.
  • Monte Carlo averages the observed returns after a state is visited.
  • The choice between TD(0) and Monte Carlo depends on whether the goal is to fit the training data or predict future data.