Concepts / n-step Methods

n-step Methods

Off-policy learning separates the policy collecting experience from the policy being learned.

  • Programming

Two Policies, One Experience

The central challenge in off-policy n-step learning is that two choices are separated. One policy generates the agent’s experience, while another policy is the one the learner wants to learn about. At the same time, the learner does not rely only on the next step: it looks ahead several steps before forming its learning information.

The behavior policy is the policy that generates experience. The target policy is the policy being learned. When these policies need not be the same, the learning process is off-policy.

producesused to learn aboutBehavior policygenerates actions andexperienceExperience trajectorystates, actions, andrewardsTarget policypolicy being learned
Which policy generates the trajectory, and which policy is being learned from that experience?

Looking Several Steps Ahead

The number n describes how far ahead the method looks before it forms its learning information. A one-step method focuses on the immediate next step. An n-step method connects the current point with events several steps later. The information from those intervening steps is part of the return, while the method can still use an estimated value farther ahead as a bootstrap.

look aheadcontinuecontinue to nthen bootstrapformsCurrent pointwhere the update beginsReward 1future informationReward 2future informationReward nlater rewardEstimated valuebootstrap from severalsteps aheadn-step returnlearning information
How do several future rewards and an estimated value from several steps ahead combine into one n-step return?

Separating the Two Questions

An agent collects a trajectory with one policy, but the learner is interested in a different stochastic policy and wants to use an n-step method.

Identify the experience source: The policy that selected the actions in the trajectory is the behavior policy.

Identify the learning target: The different policy the learner is trying to learn is the target policy.

Apply the n-step idea: The update looks ahead n steps instead of restricting its learning information to only the next step.

Address the mismatch: Because the behavior and target policies differ, the method must account for the policy mismatch. Importance sampling and tree backups provide two approaches.

This is an off-policy n-step problem: the trajectory comes from the behavior policy, the learning target is the target policy, and the update uses information from several steps ahead.

Two Ways to Resolve Mismatch

The off-policy n-step problem combines two requirements: use a trajectory produced by the behavior policy, and learn about the target policy across multiple steps. The source describes two approaches. Importance sampling changes the influence of experience through reweighting. Tree backups avoid importance sampling and instead extend the idea of Q-learning to multiple steps when the target policy is stochastic.

reweightinfluenceextend tofollowcombineObserved actionfrom behavior policyExperience weighttarget-policy likelihoodReweighted updateimportance samplingAlternative actionstarget-policy branchesMulti-step branchestree backupTree-based updateno importance sampling
How does information from alternative actions flow into an update under the two approaches?
ApproachMain ideaCentral issue or feature
Importance samplingReweight experience according to how likely it is under the target policyCan suffer high variance
Tree backupsUse a multi-step extension of Q-learning for a stochastic target policyProvide an alternative that does not use importance sampling

From n-step Prediction to Sarsa

n-step methods are not limited to prediction. When combined with Sarsa, they produce n-step Sarsa, an on-policy temporal-difference control method. The construction changes the objects being updated: instead of organizing the backup around states alone, it organizes the backup around state-action pairs.

n-step Sarsa is the n-step version of Sarsa and an on-policy temporal-difference control method.

transitionselectcontinue for n stepsexpress return withState-action pairstarting pairNext stateafter an actionNext actionε-greedy selectionLater state-actionpairn-step endpointAction-value returnestimated action values
How does n-step Sarsa move through state-action pairs under an ε-greedy policy and form an action-value return?

The structural change is easy to miss. In n-step temporal-difference prediction, backup diagrams are organized around states. In n-step Sarsa, states and actions alternate. The backup starts with an action and ends with an action, because the learned quantity is tied to a state-action pair rather than to a state by itself.

Mistakes in the Mental Model

  • Calling a method off-policy merely because it looks ahead several steps.

    The n-step property describes the look-ahead distance. Off-policy describes the separation between the behavior policy and the target policy.

    Fix: Track both dimensions separately: ask how far the method looks ahead and whether the experience-generating policy differs from the policy being learned.

  • Assuming importance sampling and tree backups solve policy mismatch in the same way.

    Importance sampling uses experience reweighting, while tree backups provide a no-importance-sampling alternative based on a multi-step extension of Q-learning.

    Fix: Associate importance sampling with reweighted experience and tree backups with multi-step Q-learning branches.

  • Ignoring the source of high variance in importance sampling.

    Across multiple steps, the reweightings combine, and the resulting update weights can become highly variable.

    Fix: Remember that multi-step reweighting can amplify variation across the trajectory.

  • Describing n-step Sarsa as a state-only method.

    n-step Sarsa replaces states with state-action pairs and alternates states and actions.

    Fix: Describe the backup as starting with an action and ending with an action, with the return expressed through estimated action values.

  • Treating ε-greedy as unrelated to the on-policy nature of n-step Sarsa.

    In n-step Sarsa, ε-greedy supplies the action-selection policy used by the on-policy control method.

    Fix: Connect ε-greedy action selection to the policy followed while learning action values.

Practice the Distinctions

MEDIUM

A learner observes a trajectory generated by one policy, wants to learn a different stochastic policy, and uses information from several steps ahead. Explain why this is off-policy n-step learning. Then describe what would change if the method were combined with Sarsa.

Hints
  • Name the policy that generated the trajectory and the policy being learned.
  • Explain what n says about the learning information.
  • For Sarsa, mention state-action pairs, estimated action values, and ε-greedy action selection.

What do you think happens?

A method has a large n. Does that guarantee that every part of the distant future has an equally strong influence on a tree-backup update?

  • Yes, because n defines the exact influence of every future branch.
  • No, the effective bootstrap can still be concentrated in a shorter portion of the future.
  • Only if the behavior and target policies are different.
Reveal answer

Answer: No, the effective bootstrap can still be concentrated in a shorter portion of the future.

A large formal look-ahead does not guarantee equally important contributions from all distant branches; the more distant contribution may become negligible.

Key Takeaways

  1. Off-policy learning separates the behavior policy that generates experience from the target policy being learned.
  2. n-step methods look ahead several steps and use information from later events rather than only the immediate next step.
  3. Importance sampling handles policy mismatch through experience reweighting, but multi-step reweighting can produce high variance.
  4. Tree backups provide a no-importance-sampling alternative based on a multi-step extension of Q-learning, although a large n can still have a short effective bootstrap.
  5. n-step Sarsa is an on-policy TD control method that uses state-action pairs, estimated action values, and an ε-greedy policy.

Key Takeaways

  • Off-policy learning uses experience from a behavior policy to learn about a separate target policy.
  • The n-step setting adds a look-ahead of n steps to the learning or planning process.
  • Importance sampling reweights experience and may have high variance, while tree backups use a multi-step Q-learning alternative without importance sampling.
  • n-step Sarsa changes the backup from states to state-action pairs and expresses the return through estimated action values.
  • ε-greedy action selection gives n-step Sarsa its policy component while preserving its on-policy TD control character.