Monte Carlo Prediction Methods
TD methods learn transition by transition, whereas Monte Carlo methods wait until an episode ends.
The Waiting-Time Difference
Monte Carlo and temporal difference methods differ in when they can learn. Monte Carlo methods wait until an episode ends, because only then is the complete return known. TD methods need to wait only one time step before learning from a transition. The central issue is therefore learning speed during an episode, not a claim that one method is always preferable.
What do you think happens?
An episode is still running, but one transition has just been observed. Which method can begin learning from that transition before the episode ends?
Reveal answer
Answer: Only the TD method
TD methods can learn after waiting one time step. Monte Carlo methods must wait until the episode ends, when the complete return is known.
Updates During an Episode
After a transition is observed, a TD method can update its value estimate after one time step. A Monte Carlo method cannot yet use the complete return for that episode, so it waits. If the episode continues for many more transitions, the difference in waiting time becomes increasingly important: TD learning may already have performed updates while Monte Carlo learning is still waiting.
Online Incremental Learning
The shorter waiting time makes TD methods naturally online and fully incremental. Online learning means that learning can begin while experience is still arriving. Incremental learning means that each newly available transition can contribute to learning rather than requiring the method to wait for an entire episode before starting its update process.
A Long Episode in Progress
Compare when TD and Monte Carlo methods can learn during an episode containing many transitions.
First transition: A TD method can begin learning after waiting one time step. A Monte Carlo method must continue waiting because the episode has not ended.
More transitions arrive: The TD method can keep learning transition by transition as new experience becomes available. The Monte Carlo method still waits for the complete return of the episode.
Episode ends: The complete return is now known, so the Monte Carlo method can learn from this episode.
TD learning begins earlier and proceeds incrementally, while Monte Carlo learning based on this episode is delayed until its end.
Very Long and Continuing Tasks
The online, incremental advantage matters especially when an episode is very long. Waiting for the complete return could delay learning for a substantial period. It also matters for continuing tasks without episodes, because there may be no natural episode ending at which a Monte Carlo method could obtain a complete return. TD methods can continue learning from transitions as they arrive instead of depending on such an endpoint.
Experimental Actions
TD methods can learn from transitions even when later actions are experimental. Once the transition caused by such an action has been observed, the TD method can use that transition after one time step. It does not need to wait for the episode's complete return before beginning to learn from the experience.
When analyzing a learning method, ask when the information needed for an update becomes available. For TD methods, the relevant opportunity arrives after one transition and one time step. For Monte Carlo methods, learning from the episode waits until its complete return is known.
Target Information
| Method | When learning can begin | Information available at that point |
|---|---|---|
| TD | After one time step | An observed transition |
| Monte Carlo | After the episode ends | The complete return |
This comparison should be kept precise. The source of TD's advantage here is the timing of the update: TD methods use each transition as it becomes available, while Monte Carlo methods postpone learning based on that episode until the complete return is known. This timing difference explains the online, incremental, long-task, and experimental-action advantages described above.
Common Timing Mistakes
Treating TD and Monte Carlo methods as if they must wait for the same event.
TD methods need to wait only one time step, whereas Monte Carlo methods wait until the episode ends because the complete return is then known.
Fix:
For each method, identify the earliest point at which its required learning information is available.Thinking that online learning means TD methods update before observing any transition.
TD methods learn from transitions, so they still need an observed transition and one time step.
Fix:
Use the more precise description: TD methods learn transition by transition as experience becomes available.Ignoring the importance of episode length.
A long wait before the episode ends can delay Monte Carlo learning while TD methods continue to update from transitions.
Fix:
Ask whether an episode is very long or whether the task has no natural episode ending.Assuming experimental actions prevent TD learning.
TD methods can learn from transitions even when later actions are experimental.
Fix:
Recognize that an observed transition can support TD learning before the episode ends.
Check Your Understanding
A task produces a very long sequence of transitions and has no natural episode ending. Explain why the timing of TD learning can be useful in this task. Then explain what would delay Monte Carlo learning based on an episode.
Hints
- Start by stating how long TD methods wait after a transition.
- Then identify when Monte Carlo methods know the complete return.
- Connect the absence of an episode ending to the availability of that return.
An experimental action produces an observed transition while the episode is still running. Which method can learn from that transition immediately after the required waiting period, and why?
Hints
- Compare one time step with an entire episode.
- Focus on the transition that has already been observed.
Key Takeaways
- TD methods learn transition by transition after waiting one time step.
- Monte Carlo methods wait until an episode ends because the complete return is then known.
- The shorter TD waiting time makes TD learning naturally online and fully incremental.
- This timing advantage matters for very long episodes and continuing tasks without episodes.
- TD methods can learn from observed transitions involving experimental actions before the episode ends.
Key Takeaways
- TD methods learn from transitions after one time step, while Monte Carlo methods wait for an episode to end.
- TD learning is naturally online and incremental because each transition can contribute as it becomes available.
- The timing advantage is especially important for very long episodes and continuing tasks without episodes.
- TD methods can continue learning from transitions produced by experimental actions before the complete return is known.