Advantages Over Dynamic Programming Methods
TD methods learn transition by transition, whereas Monte Carlo methods wait until an episode ends.
The Waiting-Time Difference
The central advantage discussed here is the timing of learning. Temporal difference methods can learn from each transition as it becomes available. Monte Carlo methods must wait until an episode ends, because only then is the complete return known. The difference may be only a matter of timing in a short episode, but it becomes important when an episode is very long or when the task continues without episodes.
One Transition at a Time
Imagine that an episode produces a sequence of transitions. TD learning can use the first transition, then the next transition, and so on as they occur. Monte Carlo learning collects the episode first and postpones learning based on that episode until its complete return is available. Therefore, TD learning begins earlier within the episode.
Comparing the Same Episode
Suppose an episode contains several transitions that become available in sequence. Compare when TD and Monte Carlo methods can learn from the first transition.
First transition appears: TD has the transition information it needs to begin learning after one time step. Monte Carlo records the transition but does not yet learn from that episode.
More transitions appear: TD can continue learning transition by transition. Monte Carlo continues waiting for the episode to finish.
Episode ends: The complete return is now known, so Monte Carlo can learn from the episode.
TD begins learning during the episode; Monte Carlo begins learning from that episode only after the episode ends.
Online Incremental Learning
The shorter waiting time makes TD methods naturally online and fully incremental. Online means that learning can begin while experience is still arriving. Incremental means that each newly available transition can contribute to learning instead of requiring the method to collect a complete episode first.
Online, incremental learning matters when an agent benefits from using information immediately rather than postponing all learning until a complete episode has finished.
Long and Continuing Tasks
The timing advantage becomes especially important for very long episodes. If an episode takes a long time to finish, a Monte Carlo method must delay learning based on that episode for the entire waiting period. TD methods can keep learning from transitions during the episode instead. The same reasoning applies to continuing tasks without episodes: when there is no terminal state at which to wait, transition-by-transition learning remains applicable.
When analyzing a learning method, ask whether the task has short episodes, very long episodes, or no episodes at all. The longer the wait for an episode to finish, or the less meaningful an episode boundary is, the more important TD's transition-by-transition timing advantage can become.
Experimental Actions Still Teach
TD methods can learn from transitions even when later actions are experimental. The transition is still newly available experience, so TD learning can use it after one time step. The action does not have to be treated as a final, already-settled choice before its resulting transition can contribute to learning.
An Exploratory Transition
An agent takes an experimental action and observes the resulting transition. Compare what TD learning can do immediately with what Monte Carlo learning must do.
Action occurs: The action is experimental, but it produces a transition that can be observed.
TD uses the transition: After one time step, TD learning can use this transition as part of its incremental learning.
Monte Carlo waits: Monte Carlo learning still waits until the episode ends before learning from that episode's complete return.
TD learning can continue learning from experience produced by experimental actions without waiting for the whole episode.
Mistakes About Timing
Assuming TD and Monte Carlo methods begin learning at the same time.
TD methods need to wait only one time step before learning from a transition, whereas Monte Carlo methods wait until the episode ends because the complete return is not known earlier.
Fix:
Track when information becomes available: TD can learn transition by transition; Monte Carlo waits for the complete episode return.Thinking that online learning means waiting for a shorter episode.
The relevant feature is that TD methods learn while transitions are arriving, making the process incremental.
Fix:
Focus on whether each new transition can contribute immediately, not merely on the eventual length of an episode.Ignoring continuing tasks because they have no episode-ending update point.
The timing advantage of TD methods applies precisely because they can learn from transitions without waiting for an episode to end.
Fix:
Use transition-by-transition learning as the key idea for continuing tasks.Assuming an experimental action produces unusable experience.
TD methods can learn from transitions even when later actions are experimental.
Fix:
Treat the resulting transition as available experience for incremental TD learning.
Check Your Understanding
A task has episodes that are so long that waiting for an episode to finish delays learning noticeably. Explain why TD methods can begin learning earlier than Monte Carlo methods. Then explain why the same timing idea is relevant to a continuing task with no episode boundary.
Hints
- State when the complete return becomes known for Monte Carlo methods.
- State how long TD methods need to wait before learning from a transition.
- Connect transition-by-transition learning to the absence of a terminal state.
What do you think happens?
An experimental action creates a transition. Which method can use that transition before the episode ends?
Reveal answer
Answer: TD methods
TD methods can learn from each transition after one time step, including transitions involving experimental actions. Monte Carlo methods wait until the episode ends so that the complete return is known.
Key Takeaways
- TD methods learn transition by transition, while Monte Carlo methods wait until an episode ends.
- TD methods need to wait only one time step before learning from a transition.
- The shorter waiting time makes TD methods naturally online and fully incremental.
- Transition-by-transition learning is especially useful for very long episodes and continuing tasks without episodes.
- TD methods can learn from transitions produced by experimental actions.
Key Takeaways
- TD methods learn from transitions as they become available; Monte Carlo methods wait for the episode's complete return.
- The timing difference makes TD learning online and incremental.
- TD methods can keep learning during very long episodes and in continuing tasks without terminal states.
- Transitions caused by experimental actions can still contribute immediately to TD learning.