Concepts / Advantages Over Dynamic Programming Methods

Advantages Over Dynamic Programming Methods

TD methods learn transition by transition, whereas Monte Carlo methods wait until an episode ends.

  • Programming

The Waiting-Time Difference

The central advantage discussed here is the timing of learning. Temporal difference methods can learn from each transition as it becomes available. Monte Carlo methods must wait until an episode ends, because only then is the complete return known. The difference may be only a matter of timing in a short episode, but it becomes important when an episode is very long or when the task continues without episodes.

One Transition at a Time

one time stepwaitcomplete returnTransitionTD receives informationTransitionMC collects informationTD updateafter one time stepEpisode endcomplete return knownMC updateafter the episode
At what point does each method receive information and update its value estimate?

Imagine that an episode produces a sequence of transitions. TD learning can use the first transition, then the next transition, and so on as they occur. Monte Carlo learning collects the episode first and postpones learning based on that episode until its complete return is available. Therefore, TD learning begins earlier within the episode.

Comparing the Same Episode

Suppose an episode contains several transitions that become available in sequence. Compare when TD and Monte Carlo methods can learn from the first transition.

First transition appears: TD has the transition information it needs to begin learning after one time step. Monte Carlo records the transition but does not yet learn from that episode.

More transitions appear: TD can continue learning transition by transition. Monte Carlo continues waiting for the episode to finish.

Episode ends: The complete return is now known, so Monte Carlo can learn from the episode.

TD begins learning during the episode; Monte Carlo begins learning from that episode only after the episode ends.

Online Incremental Learning

observelearncontinuelearn againValue estimatebefore transitionTransition 1new informationValue estimateafter transition 1Transition 2more informationValue estimateafter transition 2
How does the value estimate change after each new transition rather than after collecting a complete episode?

The shorter waiting time makes TD methods naturally online and fully incremental. Online means that learning can begin while experience is still arriving. Incremental means that each newly available transition can contribute to learning instead of requiring the method to collect a complete episode first.

Online, incremental learning matters when an agent benefits from using information immediately rather than postponing all learning until a complete episode has finished.

Long and Continuing Tasks

The timing advantage becomes especially important for very long episodes. If an episode takes a long time to finish, a Monte Carlo method must delay learning based on that episode for the entire waiting period. TD methods can keep learning from transitions during the episode instead. The same reasoning applies to continuing tasks without episodes: when there is no terminal state at which to wait, transition-by-transition learning remains applicable.

next transitionnext transitioneventual endpointTransition 1learning can beginTransition 2learning continuesTransition 3more learningTerminal statemay be very far away orabsent
How can TD methods keep learning when an episode is too long to wait for or has no terminal state?

When analyzing a learning method, ask whether the task has short episodes, very long episodes, or no episodes at all. The longer the wait for an episode to finish, or the less meaningful an episode boundary is, the more important TD's transition-by-transition timing advantage can become.

Experimental Actions Still Teach

causescontributes toExperimental actionaction is takenObserved transitionnew experienceTD learninguse transition after onetime step
How does a transition caused by an exploratory action immediately contribute to learning?

TD methods can learn from transitions even when later actions are experimental. The transition is still newly available experience, so TD learning can use it after one time step. The action does not have to be treated as a final, already-settled choice before its resulting transition can contribute to learning.

An Exploratory Transition

An agent takes an experimental action and observes the resulting transition. Compare what TD learning can do immediately with what Monte Carlo learning must do.

Action occurs: The action is experimental, but it produces a transition that can be observed.

TD uses the transition: After one time step, TD learning can use this transition as part of its incremental learning.

Monte Carlo waits: Monte Carlo learning still waits until the episode ends before learning from that episode's complete return.

TD learning can continue learning from experience produced by experimental actions without waiting for the whole episode.

Mistakes About Timing

  • Assuming TD and Monte Carlo methods begin learning at the same time.

    TD methods need to wait only one time step before learning from a transition, whereas Monte Carlo methods wait until the episode ends because the complete return is not known earlier.

    Fix: Track when information becomes available: TD can learn transition by transition; Monte Carlo waits for the complete episode return.

  • Thinking that online learning means waiting for a shorter episode.

    The relevant feature is that TD methods learn while transitions are arriving, making the process incremental.

    Fix: Focus on whether each new transition can contribute immediately, not merely on the eventual length of an episode.

  • Ignoring continuing tasks because they have no episode-ending update point.

    The timing advantage of TD methods applies precisely because they can learn from transitions without waiting for an episode to end.

    Fix: Use transition-by-transition learning as the key idea for continuing tasks.

  • Assuming an experimental action produces unusable experience.

    TD methods can learn from transitions even when later actions are experimental.

    Fix: Treat the resulting transition as available experience for incremental TD learning.

Check Your Understanding

MEDIUM

A task has episodes that are so long that waiting for an episode to finish delays learning noticeably. Explain why TD methods can begin learning earlier than Monte Carlo methods. Then explain why the same timing idea is relevant to a continuing task with no episode boundary.

Hints
  • State when the complete return becomes known for Monte Carlo methods.
  • State how long TD methods need to wait before learning from a transition.
  • Connect transition-by-transition learning to the absence of a terminal state.

What do you think happens?

An experimental action creates a transition. Which method can use that transition before the episode ends?

  • TD methods
  • Monte Carlo methods
  • Neither method
Reveal answer

Answer: TD methods

TD methods can learn from each transition after one time step, including transitions involving experimental actions. Monte Carlo methods wait until the episode ends so that the complete return is known.

Key Takeaways

  1. TD methods learn transition by transition, while Monte Carlo methods wait until an episode ends.
  2. TD methods need to wait only one time step before learning from a transition.
  3. The shorter waiting time makes TD methods naturally online and fully incremental.
  4. Transition-by-transition learning is especially useful for very long episodes and continuing tasks without episodes.
  5. TD methods can learn from transitions produced by experimental actions.

Key Takeaways

  • TD methods learn from transitions as they become available; Monte Carlo methods wait for the episode's complete return.
  • The timing difference makes TD learning online and incremental.
  • TD methods can keep learning during very long episodes and in continuing tasks without terminal states.
  • Transitions caused by experimental actions can still contribute immediately to TD learning.