Off-Policy Methods
Function approximation introduces issues including nonstationarity, bootstrapping, and delayed targets.
Two Questions About Experience
When an agent explores, its experience can be examined from two different viewpoints. First, which policy is actually generating the agent's behavior? Second, which policy's value function is the learning process trying to evaluate? Off-policy methods are built around the possibility that these two policies are different.
The word off-policy does not mean that the agent has stopped following a policy. It means that the policy being followed and the policy whose value is being learned may not be the same.
Following Versus Evaluating
On-policy learning evaluates the policy currently being followed. The behavior policy and the policy whose value is learned are therefore treated as the same policy.
Off-policy learning evaluates the policy currently considered best, even when that policy differs from the behavior policy that generated the experience.
Identifying the Two Policies
An agent collects experience while exploring, but the learning objective is to evaluate the policy currently considered best. Is this on-policy or off-policy learning?
Identify the behavior policy: The behavior policy is the policy the agent is actually following while it collects experience.
Identify the target policy: The target policy is the policy whose value is being evaluated. Here, it is the policy currently considered best.
Compare the roles: Because exploration can make the behavior policy differ from the policy currently considered best, the two policies may not be identical.
This is an off-policy setting when the policy being followed differs from the policy whose value is being learned.
Exploration Splits the Policies
Exploration is the reason the two policy roles can separate. An agent may follow a behavior policy that collects varied experience, while the learning objective concerns a different policy, such as the policy currently considered best. In that case, the actions producing the data and the policy being evaluated are not described by the same role.
Imagine separating a learner's data-collection role from its evaluation goal. The data-collection role explores, while the evaluation goal asks about the policy currently considered best. The important point is not the particular exploratory action; it is that the policy generating experience need not be the policy whose value is learned.
What do you think happens?
An agent follows one policy to collect experience but learns the value of another policy. Which classification fits?
Reveal answer
Answer: Off-policy
Off-policy learning concerns the case where the policy being followed can differ from the policy whose value is being learned.
Function Approximation Warning Signs
The on-policy and off-policy distinction is only one way to organize reinforcement-learning methods. The source also asks learners to examine this distinction alongside bootstrapping and function approximation. These dimensions should be considered together rather than treating every method as an isolated case.
Function approximation does more than add a new representation to ordinary reinforcement learning. The chapter overview highlights problems that are not normally encountered in conventional supervised learning. Three warning signs are named: nonstationarity, bootstrapping, and delayed targets.
- Nonstationarity is one of the difficulties that must be considered when reinforcement learning uses function approximation.
- Bootstrapping is another dimension alongside the on-policy or off-policy distinction.
- Delayed targets are a third issue highlighted in the overview.
- These terms should initially be treated as warning signs rather than as a list of detailed algorithms to memorize.
The Chapter Roadmap
The chapter sequence progresses from a restricted training setting toward broader and alternative approaches. The early chapters use on-policy training. First, prediction studies a given policy and approximates its value function. Then control aims to find an approximation to the optimal policy. After that, off-policy methods broaden the scope by separating the policy being followed from the policy whose value is learned.
| Chapter focus | Main learning role | Question to ask |
|---|---|---|
| Prediction | Approximate the value function for a given policy | What is the value of this policy? |
| Control | Find an approximation to the optimal policy | How can learning move toward an optimal policy? |
| Off-policy methods | Study learning when the followed policy and evaluated policy may differ | Whose value function is being learned? |
| Eligibility traces | Analyze a computational mechanism for multistep methods | How can eligibility traces affect the computational properties of multistep methods? |
| Policy-gradient methods | Approximate the optimal policy directly | Can the policy be approximated directly without first forming an approximate value function? |
A roadmap for matching each chapter with its learning objective.
The roadmap separates learning settings from learning mechanisms. On-policy and off-policy describe training settings, while eligibility traces describe an algorithmic mechanism and policy-gradient methods describe a different approach to control.
Reading Methods by Dimensions
Do not classify a reinforcement-learning method using only one label. Begin by asking whether the policy being followed is also the policy whose value is learned. Then examine the method alongside bootstrapping and function approximation. This combined view helps organize methods by their properties and prepares you to understand why function approximation requires the learning problem to be reconsidered.
Assuming that exploration automatically makes a method off-policy.
Exploration can make the policies different, but the classification depends on whether the policy being evaluated is the same as the policy being followed.
Fix:
Identify both policy roles before assigning the label.Defining off-policy learning as learning without a policy.
Off-policy learning still involves a policy that generates experience.
Fix:
Separate the behavior policy from the target policy whose value is learned.Treating prediction and control as the same chapter objective.
The roadmap assigns prediction to the case where a policy is given and its value function is approximated. Control then aims toward an optimal policy.
Fix:
Use prediction for evaluating a given policy and control for the objective of finding an approximation to the optimal policy.Treating function approximation as only a change of representation.
The overview highlights nonstationarity, bootstrapping, and delayed targets as additional issues.
Fix:
Treat those three terms as warning signs when analyzing reinforcement learning with function approximation.Confusing eligibility traces with a separate policy category.
Eligibility traces are introduced as an algorithmic mechanism with a computational role in multistep methods.
Fix:
Keep training settings, such as on-policy or off-policy, distinct from mechanisms such as eligibility traces.
For each situation, identify the policy being followed, the policy whose value is learned, and whether the setting is on-policy or off-policy. Then state which roadmap topic is most relevant: prediction, control, off-policy methods, eligibility traces, or policy-gradient methods.
Hints
- Start by identifying the behavior policy.
- Next identify the target policy.
- If the two policy roles differ, classify the setting as off-policy.
- Use prediction when a given policy is being evaluated, and control when the aim is an approximation to the optimal policy.
- Use eligibility traces for the mechanism and computational role in multistep methods; use policy-gradient methods for direct approximation of the optimal policy.
Key Takeaways
- On-policy learning evaluates the policy currently being followed.
- Off-policy learning evaluates a policy that may differ from the behavior policy, especially when exploration separates the two roles.
- Function approximation brings three highlighted warning signs: nonstationarity, bootstrapping, and delayed targets.
- The roadmap moves from prediction with a given policy to control, off-policy methods, eligibility traces, and finally direct policy-gradient control.
- The most useful classification question is whose value function is being learned, considered alongside bootstrapping and function approximation.
Key Takeaways
- On-policy and off-policy methods differ according to whether the policy being followed is also the policy whose value is learned.
- Exploration can create separate behavior and target policies, making off-policy analysis necessary.
- Function approximation requires attention to nonstationarity, bootstrapping, and delayed targets.
- The chapter sequence progresses from prediction and control to off-policy methods, eligibility traces, and policy-gradient methods.