Concepts / Temporal-Difference Learning Methods

Temporal-Difference Learning Methods

Off-policy learning separates the policy collecting experience from the policy being learned.

  • Programming

Two Policies in One Learning Process

Temporal-difference learning methods address a central problem in finite Markov decision problems: how can an agent improve its estimates while experience is being generated? The distinctive setting here is off-policy learning. One policy collects the experience, while another policy is the one being learned.

The behavior policy is the policy that generates experience. The target policy is the policy being learned. When these policies are different, the learning process is off-policy.

generatessuppliesdefinesBehavior policygenerates experienceExperiencetrajectoryTarget policypolicy being learnedLearning updateuses experience for target
How does experience collected by one policy flow into learning a different target policy?

The policies need not be the same. The behavior policy determines what the agent actually experiences, while the target policy determines what the learner wants to evaluate or improve. The policy mismatch is therefore not an incidental detail: it is the defining feature that makes the learning off-policy.

Generated example: imagine an agent collecting a trajectory under one stochastic policy while the learner is interested in a different stochastic policy. The learner must use the collected trajectory while accounting for the fact that it was produced by the behavior policy rather than the target policy.

Looking Several Steps Ahead

An n-step method does not restrict its learning information to only the next step. The number n describes how far ahead the method looks before forming that information. As n grows, the method connects the current point with events farther along the trajectory or plan.

advanceadvancecontinue to ncontributesCurrent pointstarting pointStep t+1one step aheadStep t+2two steps aheadStep t+nn steps aheadLearning informationlook-ahead of n steps
What information from each step ahead is combined with the bootstrap estimate, and how does the return change as n increases?

The n-step idea applies both to learning and to planning. In either use, the method adds a look-ahead of n steps instead of relying only on the immediate next step. In the off-policy setting, this longer look-ahead must be combined with a way to handle the difference between the behavior and target policies.

Changing the Look-Ahead

A learner starts from the current point and considers an n-step method. What changes when n is one, two, or several steps?

n equals one: The learning information is based on the next step, so the method uses the shortest look-ahead described here.

n equals two: The method connects the current point with information reaching two steps ahead rather than stopping after the first step.

n is larger: The method incorporates events farther along the trajectory or plan. The central change is the distance of the look-ahead, not a change in the definition of behavior and target policies.

Increasing n means looking farther ahead before forming the learning information.

Two Ways to Correct the Policy Mismatch

Once an n-step method uses experience from the behavior policy, it must address the fact that the target policy may be different. The source describes two approaches: importance sampling and tree backups.

handled byhandled byproducesextends acrossImportance samplingreweights experienceTree backupsmulti-step Q-learningPolicy mismatchbehavior versus targetReweighted experiencetarget-policy likelihoodBranchingpossibilitiesstochastic target policy
How do importance sampling and tree backups use the same branching possibilities differently to estimate the target policy's value?

Importance sampling handles the mismatch through experience reweighting. The influence of an observed experience is changed according to how likely that experience is under the target policy. Tree backups take a different route. They extend the idea of Q-learning to multiple steps when the target policy is stochastic, using the branching possibilities rather than relying on importance sampling.

MethodMain response to policy mismatchSource-stated concern or feature
Importance samplingReweights experience according to target-policy likelihoodMay suffer high variance
Tree backupsUses a multi-step extension of Q-learning for a stochastic target policyAvoids importance sampling

The two approaches address the same off-policy problem in different ways.

Why Reweighting Can Become Unstable

Importance sampling can become high variance because it reweights experience to account for the target policy. Across an n-step look-ahead, the adjustment concerns several steps of experience. As more stepwise likelihood adjustments accumulate, the resulting influence can vary greatly from one trajectory to another.

contributescontributescontributescan produceStep 1 ratiotarget versus behaviorStep 2 ratiotarget versus behaviorFurther ratiosacross n stepsCombined weightexperience influenceHigh variancepossible outcome
How do products of action-probability ratios accumulate across several steps and cause high variance?

The important practical lesson is not that importance sampling fails, but that its correction can be noisy. A long n-step sequence creates more opportunities for the reweighting to differ across experiences. That is why the source identifies high variance as a weakness of the approach.

Long Horizons and Short Effective Bootstraps

Tree backups can use an n-step horizon while still having only a short effective bootstrap. The distinction is between the formal distance n and the part of the backup that contributes most strongly to the resulting estimate. A large n says that the method can look far ahead; it does not say that every step in that horizon has equal practical influence.

branchescontinues to ncontributes stronglyalso contributesCurrent statebackup beginsNear branchesshort-range contributionFar brancheswithin n-step horizonBackup estimatecombined result
How can a tree backup use a long n-step horizon while most of its contribution comes from only a few steps ahead?

This is why n and effective bootstrap should not be treated as synonyms. The tree backup can represent a multi-step extension over a long horizon, yet the update's effective bootstrap may remain concentrated near the current point. The source presents this as an important limitation to recognize when interpreting what a large n accomplishes.

When comparing n-step methods, ask two separate questions: how far can the method look ahead, and how much of the final estimate is effectively determined by the nearer part of the backup? Keeping these questions separate prevents a large n from being mistaken for uniformly long-range influence.

Three Method Classes

Temporal-difference learning is one of three broad method classes considered here for solving finite Markov decision problems. The most useful comparison asks what information a method requires and when it can make progress.

requiresdoes not requiredoes not requiresupportsDynamic programmingcomplete accurate modelMonte Carlono modelTemporal differenceno model, incrementalModel requirementcomplete and accurateIncrementalcomputationstep-by-step progress
What are the differences in model requirements, update timing, bootstrapping, and incrementality among the three method classes?
Method classModel requirementIncremental computationMain stated strength or weakness
Dynamic programmingRequires a complete and accurate modelNot identified here as the defining featureMathematically well developed
Monte CarloDoes not require a modelNot well suited to step-by-step incremental computationConceptually simple
Temporal-difference learningDoes not require a modelFully incrementalFlexible, but more complex to analyze

The comparison focuses on model access and the timing of computation.

In this comparison, requiring no model means that the method does not depend on being supplied with a complete and accurate model of the environment. Fully incremental computation means that learning can make progress step by step rather than waiting for a complete outcome before progressing.

Common Comparison Errors

  • Calling a method off-policy merely because it uses several steps of experience.

    Off-policy learning is defined by a difference between the behavior policy and the target policy.

    Fix: Check whether the policy generating experience differs from the policy being learned.

  • Treating n as the amount of effective influence in every backup.

    The source distinguishes a long n-step horizon from a short effective bootstrap in tree backups.

    Fix: Separate the formal look-ahead distance from where the backup's effective contribution is concentrated.

  • Assuming importance sampling removes all uncertainty.

    Importance sampling handles policy mismatch through reweighting but may suffer high variance.

    Fix: Recognize policy correction and variance as separate issues.

  • Saying that all model-free methods are fully incremental.

    The source says Monte Carlo methods are not well suited to step-by-step incremental computation, while temporal-difference methods are fully incremental.

    Fix: Compare model requirements and computation timing separately.

  • Claiming that temporal-difference learning is always the fastest or best method.

    The source identifies trade-offs and does not give a universal ranking for efficiency or convergence speed.

    Fix: State its specific advantage: no model requirement combined with incremental computation.

Check Your Understanding

MEDIUM

A behavior policy generates a trajectory, and a different stochastic target policy is being learned with an n-step method. Explain why the setting is off-policy, what n contributes, and which two approaches can address the policy mismatch.

Hints
  • Identify the policy that generates experience and the policy being learned.
  • Describe n as a look-ahead distance.
  • Name reweighting for importance sampling and the multi-step Q-learning extension for tree backups.
EASY

Compare the following choices for solving a finite Markov decision problem: dynamic programming, Monte Carlo methods, and temporal-difference learning. For each, state whether it requires a complete and accurate model and whether it is fully incremental.

Hints
  • Dynamic programming requires a complete and accurate model.
  • Monte Carlo methods do not require a model but are not well suited to step-by-step incremental computation.
  • Temporal-difference methods require no model and are fully incremental.

What do you think happens?

A learner increases n in an off-policy method. Does that automatically eliminate the policy mismatch?

  • Yes, because a longer look-ahead makes the policies equivalent
  • No, because look-ahead distance and policy mismatch are separate issues
Reveal answer

Answer: No, because look-ahead distance and policy mismatch are separate issues.

n controls how far ahead the method looks. Off-policy learning concerns the difference between the behavior policy and the target policy, which still requires an approach such as importance sampling or tree backups.

Key Takeaways

  1. Off-policy learning uses experience generated by a behavior policy to learn about a different target policy.
  2. An n-step method looks ahead n steps, connecting the current point with information farther along a trajectory or plan.
  3. Importance sampling handles policy mismatch through experience reweighting but may have high variance.
  4. Tree backups avoid importance sampling by extending Q-learning to multiple steps for a stochastic target policy; a long n does not guarantee a long effective bootstrap.
  5. Dynamic programming requires a complete and accurate model, Monte Carlo methods need no model but are not well suited to step-by-step incremental computation, and temporal-difference learning needs no model while remaining fully incremental.

Key Takeaways

  • Off-policy learning separates the policy collecting experience from the policy being learned.
  • n-step methods extend the look-ahead beyond the immediate next step.
  • Importance sampling reweights experience and can suffer high variance, while tree backups provide a no-importance-sampling alternative.
  • Tree backups may have a short effective bootstrap even when their formal n-step horizon is large.
  • Temporal-difference learning requires no model and is fully incremental, although it is more complex to analyze.