Returns and Monte Carlo Prediction
TD(0) and Monte Carlo can disagree because they use different evidence from the same episodes.
Introduction to Returns and Monte Carlo Prediction
TD(0) and Monte Carlo methods are used to estimate the value of states in a Markov reward process. Although they use the same episodes, they can produce different estimates for the same state. This difference arises because they utilize different evidence from the episodes.
How TD(0) and Monte Carlo Assign Different Values
The key difference between TD(0) and Monte Carlo lies in how they estimate the value of a state. Monte Carlo follows the complete observed returns after visiting a state, whereas TD(0) uses the estimated value of the next state it transitions to.
Batch TD(0) Bootstrapping
Batch TD(0) updates the estimate of a state's value by using the estimated value of the next state it transitions to. This process is called bootstrapping.
Monte Carlo Estimation
Monte Carlo estimation involves averaging the returns observed after visiting a state. The value of a state is the average return following visits to that state.
Fitting Training Data vs Predicting Future Data
Monte Carlo is good at fitting the existing training data, as it directly averages the observed returns. However, TD(0) is better at predicting future data because it uses the Markov property to generalize from the transition patterns observed.
The Markov Property and TD(0)
The Markov property states that the future depends only on the current state, not on the history leading to it. TD(0) leverages this property by using the next state's value to update the current state's value, making it effective for prediction in Markov processes.
Common Mistakes
Assuming TD(0) and Monte Carlo always give the same estimates.
TD(0) uses bootstrapping from the next state, while Monte Carlo averages observed returns.
Fix:
Understand that different methods use different evidence from the same episodes.Thinking Monte Carlo is always better because it directly uses observed returns.
While Monte Carlo fits the training data well, it may not predict future data as effectively as TD(0) in Markov processes.
Fix:
Consider the Markov property and the need for generalization when choosing between methods.
Practice
Consider a simple Markov reward process with two states, A and B, where A transitions to B with a reward of 0, and B terminates with a reward of 1 or 0. Use both TD(0) and Monte Carlo methods to estimate the value of state A after observing several episodes.
Hints
- Start by calculating the observed returns for state A.
- Use the estimated value of B to update the value of A in TD(0).
Summary
- TD(0) and Monte Carlo can assign different values to the same state due to their different approaches to estimation.
- TD(0) bootstraps from the estimated value of the next state.
- Monte Carlo estimates a state's value by averaging the observed returns after visiting it.
- The Markov property favors TD(0) for prediction in Markov processes.
Key Takeaways
- TD(0) and Monte Carlo methods can produce different estimates for the same state due to their different uses of episode data.
- TD(0) uses bootstrapping from the next state's estimated value.
- Monte Carlo averages the observed returns after a state is visited.
- The choice between TD(0) and Monte Carlo depends on whether the goal is to fit the training data or predict future data.