Multi-step reinforcement learning
Off-policy learning separates the policy that generates experience from the policy being learned.
One Experience, Two Policies
A reinforcement-learning agent can learn from experience generated by a policy different from the policy it is trying to learn. The policy that generates behavior is called the behavior policy. The policy being learned is called the target policy. This separation is the central idea of off-policy learning.
The experience and the learning objective do not have to come from the same policy. The behavior policy supplies what happened; the target policy determines how that experience contributes to learning.
Following the Sampled Path
The n-step tree backup algorithm is an off-policy method. It follows an observed action through its sampled next state, preserving the part of the experience that actually occurred. After reaching that next state, the backup can expand over possible actions instead of waiting for every one of them to be sampled.
A short backup trace
Suppose the behavior policy samples an action from state s0 and the environment produces state s1. At s1, the target policy assigns probabilities to three possible actions.
Step 1: retain the observation: The backup uses the observed action from s0 and the sampled transition to s1. This is the sample-transition part of the tree backup.
Step 2: expand at s1: Rather than using only the action that the behavior policy happened to sample next, the backup considers all actions available at s1.
Step 3: weight the branches: Each action branch contributes according to the probability assigned to that action by the target policy.
The backup is neither a purely sampled path nor a full expansion from the beginning. It combines an observed path segment with target-policy-weighted action branches.
Adding Unsampled Actions
A sampled transition tells the algorithm what happened for one action. It does not, by itself, tell the algorithm to ignore the other actions. In a full backup, the other actions enter through their possible contributions, weighted by the probabilities that the target policy assigns to them.
Stopping the Backup
The n-step tree backup does not expand the tree indefinitely. It stops when the chosen n-step backup depth has been reached. If a terminal state is reached before that depth, there is no later state or action branch to continue backing up, so the remaining part of the return ends there.
Three Research Phases
Research on teams of reinforcement-learning agents developed through three broad phases. The phases are best understood as a widening sequence of questions: first, how learning automata could learn in bandit, team, and game problems; second, how learning could use context and connect with artificial neural networks; and third, how learning rules could reflect findings about the brain.
- First phase: research associated with M. L. Tsetlin and later stochastic learning automata studied bandit, team, and game problems.
- Second phase: researchers extended learning automata to associative or contextual learning and connected them with single-layer and multi-layer artificial neural networks.
- Third phase: research paid greater attention to synaptic plasticity and constraints suggested by neuroscience.
From Non-Associative to Contextual
| Approach | What learning uses | Historical role |
|---|---|---|
| Non-associative learning automata | Learning behavior in bandit, team, and game problems without an explicitly associative context | Foundation of the first research phase |
| Associative or contextual reinforcement learning | Context together with reinforcement | Second phase; connected learning automata with artificial neural networks |
| Associative reward-penalty algorithm | Associative reinforcement learning in learning units and teams | Strengthened connections among learning automata, pattern classification, associative reinforcement learning, and artificial neural networks |
The word non-associative marks the boundary of the first phase: those studies investigated learning behavior but did not address the contextual bandit case. Associative reinforcement learning added an explicit context, so the learner could connect what it perceived with the action and the reinforcement received.
The associative reward-penalty algorithm, introduced by Barto and Anandan in 1985, was important because it connected several lines of work: stochastic learning automata, pattern classification, associative reinforcement learning, and artificial neural networks. Teams of these units were connected into multi-layer neural networks and were reported to learn nonlinear functions, including XOR, while receiving a globally broadcast reinforcement signal. Later mathematical analysis showed that a special case of A R - P is a REINFORCE algorithm.
Neuroscience and the Third Phase
The third phase was influenced by growing neuroscience support for the possibility that this kind of learning occurs in the brain. Researchers therefore paid more attention to synaptic plasticity and to biological constraints, rather than treating the learning rule only as an abstract computational procedure.
- STDP
- Dopamine
- Reward-modulated STDP
- Synaptic plasticity
- Other biological constraints suggested by neuroscience
Mistakes to Avoid
Treating the behavior policy and target policy as the same thing.
Off-policy learning separates the policy generating experience from the policy being learned.
Fix:
Use the behavior policy to identify the observed experience and the target policy to determine the learning update and branch weights.Following only the sampled action at every step.
The n-step tree backup combines sampled transitions with full backups over possible actions.
Fix:
Follow the observed transition where sampling is used, then include alternative actions through the full backup.Giving every alternative action equal influence.
Target-policy probabilities determine the weights of state-to-action branches.
Fix:
Weight each branch according to the probability assigned by the target policy.Calling the first phase associative simply because it involved learning.
The first phase did not address the contextual bandit case.
Fix:
Reserve associative or contextual learning for approaches that explicitly use context.Reducing the importance of A R - P to a single learning unit.
Its historical importance came from linking several research traditions and supporting teams of units in multi-layer networks.
Fix:
Remember both its associative learning role and its connections to neural networks and REINFORCE.
Check Your Understanding
An agent collects a transition using one policy, but the learning algorithm is updating a different target policy. Explain what information comes from the behavior policy, what information comes from the target policy, and why the backup includes actions that were not sampled.
Hints
- Start by separating experience generation from policy evaluation or improvement.
- Identify the observed transition as the sampled part of the backup.
- Explain that alternative actions enter through a full backup weighted by target-policy probabilities.
Place these developments in order and state what each added: non-associative learning automata, associative or contextual reinforcement learning, and neuroscience-informed learning rules.
Hints
- The first phase focused on bandit, team, and game problems.
- The second phase added context and connections with artificial neural networks.
- The third phase added attention to synaptic plasticity and other biological constraints.
Key Takeaways
- Off-policy learning lets an agent learn a target policy from experience generated by a different behavior policy.
- n-step tree backup alternates between sampled transitions and full backups over possible actions.
- Alternative actions contribute through branches weighted by target-policy probabilities.
- The backup stops when its n-step depth is reached or when a terminal state ends the return.
- Research on teams of reinforcement-learning agents moved from learning automata, to associative neural learning, and then toward neuroscience-informed learning rules.
- A R - P was important because it connected associative reinforcement learning with learning automata, pattern classification, artificial neural networks, and later REINFORCE analysis.
Key Takeaways
- Off-policy learning separates the policy that generates experience from the policy being learned.
- n-step tree backup follows observed transitions while also performing full, target-policy-weighted backups over alternative actions.
- The process stops at the selected backup depth or when a terminal state ends further expansion.
- Research on teams of reinforcement-learning agents developed through learning automata, associative neural learning, and neuroscience-informed phases.
- The associative reward-penalty algorithm helped connect these traditions and was later related to REINFORCE.