Concepts / Multi-step reinforcement learning

Multi-step reinforcement learning

Off-policy learning separates the policy that generates experience from the policy being learned.

  • Programming

One Experience, Two Policies

A reinforcement-learning agent can learn from experience generated by a policy different from the policy it is trying to learn. The policy that generates behavior is called the behavior policy. The policy being learned is called the target policy. This separation is the central idea of off-policy learning.

producesupdatesprovidesguides backupBehavior policygenerates experienceObserved transitionssampled actions and nextstatesTarget policypolicy being learnedTarget probabilitiesweights for action branches
How can experience generated by one policy be used to learn a different policy?

The experience and the learning objective do not have to come from the same policy. The behavior policy supplies what happened; the target policy determines how that experience contributes to learning.

Following the Sampled Path

The n-step tree backup algorithm is an off-policy method. It follows an observed action through its sampled next state, preserving the part of the experience that actually occurred. After reaching that next state, the backup can expand over possible actions instead of waiting for every one of them to be sampled.

sampleobserved transitionfull backupfull backupfull backupState s0Observed actionsampledState s1sampled next stateAction a1alternative branchAction a2alternative branchAction a3alternative branch
What happens as the backup follows sampled transitions for some steps and then expands all possible actions?

A short backup trace

Suppose the behavior policy samples an action from state s0 and the environment produces state s1. At s1, the target policy assigns probabilities to three possible actions.

Step 1: retain the observation: The backup uses the observed action from s0 and the sampled transition to s1. This is the sample-transition part of the tree backup.

Step 2: expand at s1: Rather than using only the action that the behavior policy happened to sample next, the backup considers all actions available at s1.

Step 3: weight the branches: Each action branch contributes according to the probability assigned to that action by the target policy.

The backup is neither a purely sampled path nor a full expansion from the beginning. It combines an observed path segment with target-policy-weighted action branches.

Adding Unsampled Actions

A sampled transition tells the algorithm what happened for one action. It does not, by itself, tell the algorithm to ignore the other actions. In a full backup, the other actions enter through their possible contributions, weighted by the probabilities that the target policy assigns to them.

observedexpandedexpandedsampled transitiontarget probabilitytarget probabilityState sSampled actionobservedSampled returnOther actiontarget probabilityAction returnweightedOther actiontarget probabilityAction returnweighted
How do actions that were not sampled still contribute to the update?

Stopping the Backup

The n-step tree backup does not expand the tree indefinitely. It stops when the chosen n-step backup depth has been reached. If a terminal state is reached before that depth, there is no later state or action branch to continue backing up, so the remaining part of the return ends there.

inspectnot reachedterminal reachedrepeatstopBackup stepStopping conditiondepth or terminal stateNext backupsample or expandBackup completeTerminal statereturn ends
What condition causes the backup process to stop, and how does reaching a terminal state change the remaining return?

Three Research Phases

Research on teams of reinforcement-learning agents developed through three broad phases. The phases are best understood as a widening sequence of questions: first, how learning automata could learn in bandit, team, and game problems; second, how learning could use context and connect with artificial neural networks; and third, how learning rules could reflect findings about the brain.

adds contextadds biological constraintsLearning automatabandit, team, game problemsAssociative learningcontext and neural networksNeuroscienceconstraintsbrain-related learningrules
How did research on teams of reinforcement-learning agents develop across three phases?
  • First phase: research associated with M. L. Tsetlin and later stochastic learning automata studied bandit, team, and game problems.
  • Second phase: researchers extended learning automata to associative or contextual learning and connected them with single-layer and multi-layer artificial neural networks.
  • Third phase: research paid greater attention to synaptic plasticity and constraints suggested by neuroscience.

From Non-Associative to Contextual

ApproachWhat learning usesHistorical role
Non-associative learning automataLearning behavior in bandit, team, and game problems without an explicitly associative contextFoundation of the first research phase
Associative or contextual reinforcement learningContext together with reinforcementSecond phase; connected learning automata with artificial neural networks
Associative reward-penalty algorithmAssociative reinforcement learning in learning units and teamsStrengthened connections among learning automata, pattern classification, associative reinforcement learning, and artificial neural networks
reinforcementconditionsreceivessupportsActionreward signalLearning behaviorContextstate informationContextual learningreward signalA R - Passociative algorithmActionchosen in context
What is the difference between learning from an action's reward alone and learning conditioned on contextual states?

The word non-associative marks the boundary of the first phase: those studies investigated learning behavior but did not address the contextual bandit case. Associative reinforcement learning added an explicit context, so the learner could connect what it perceived with the action and the reinforcement received.

The associative reward-penalty algorithm, introduced by Barto and Anandan in 1985, was important because it connected several lines of work: stochastic learning automata, pattern classification, associative reinforcement learning, and artificial neural networks. Teams of these units were connected into multi-layer neural networks and were reported to learn nonlinear functions, including XOR, while receiving a globally broadcast reinforcement signal. Later mathematical analysis showed that a special case of A R - P is a REINFORCE algorithm.

Neuroscience and the Third Phase

The third phase was influenced by growing neuroscience support for the possibility that this kind of learning occurs in the brain. Researchers therefore paid more attention to synaptic plasticity and to biological constraints, rather than treating the learning rule only as an abstract computational procedure.

  • STDP
  • Dopamine
  • Reward-modulated STDP
  • Synaptic plasticity
  • Other biological constraints suggested by neuroscience

Mistakes to Avoid

  • Treating the behavior policy and target policy as the same thing.

    Off-policy learning separates the policy generating experience from the policy being learned.

    Fix: Use the behavior policy to identify the observed experience and the target policy to determine the learning update and branch weights.

  • Following only the sampled action at every step.

    The n-step tree backup combines sampled transitions with full backups over possible actions.

    Fix: Follow the observed transition where sampling is used, then include alternative actions through the full backup.

  • Giving every alternative action equal influence.

    Target-policy probabilities determine the weights of state-to-action branches.

    Fix: Weight each branch according to the probability assigned by the target policy.

  • Calling the first phase associative simply because it involved learning.

    The first phase did not address the contextual bandit case.

    Fix: Reserve associative or contextual learning for approaches that explicitly use context.

  • Reducing the importance of A R - P to a single learning unit.

    Its historical importance came from linking several research traditions and supporting teams of units in multi-layer networks.

    Fix: Remember both its associative learning role and its connections to neural networks and REINFORCE.

Check Your Understanding

MEDIUM

An agent collects a transition using one policy, but the learning algorithm is updating a different target policy. Explain what information comes from the behavior policy, what information comes from the target policy, and why the backup includes actions that were not sampled.

Hints
  • Start by separating experience generation from policy evaluation or improvement.
  • Identify the observed transition as the sampled part of the backup.
  • Explain that alternative actions enter through a full backup weighted by target-policy probabilities.
EASY

Place these developments in order and state what each added: non-associative learning automata, associative or contextual reinforcement learning, and neuroscience-informed learning rules.

Hints
  • The first phase focused on bandit, team, and game problems.
  • The second phase added context and connections with artificial neural networks.
  • The third phase added attention to synaptic plasticity and other biological constraints.

Key Takeaways

  1. Off-policy learning lets an agent learn a target policy from experience generated by a different behavior policy.
  2. n-step tree backup alternates between sampled transitions and full backups over possible actions.
  3. Alternative actions contribute through branches weighted by target-policy probabilities.
  4. The backup stops when its n-step depth is reached or when a terminal state ends the return.
  5. Research on teams of reinforcement-learning agents moved from learning automata, to associative neural learning, and then toward neuroscience-informed learning rules.
  6. A R - P was important because it connected associative reinforcement learning with learning automata, pattern classification, artificial neural networks, and later REINFORCE analysis.

Key Takeaways

  • Off-policy learning separates the policy that generates experience from the policy being learned.
  • n-step tree backup follows observed transitions while also performing full, target-policy-weighted backups over alternative actions.
  • The process stops at the selected backup depth or when a terminal state ends further expansion.
  • Research on teams of reinforcement-learning agents developed through learning automata, associative neural learning, and neuroscience-informed phases.
  • The associative reward-penalty algorithm helped connect these traditions and was later related to REINFORCE.