On-Policy Prediction and Control
Function approximation introduces issues including nonstationarity, bootstrapping, and delayed targets.
Why the Roadmap Matters
Function approximation changes more than the representation used in reinforcement learning. The chapter overview emphasizes that it also introduces problems that are not normally encountered in conventional supervised learning. The three warning signs are nonstationarity, bootstrapping, and delayed targets. The chapter sequence is organized to introduce these issues progressively and to reconsider learning objectives in light of them.
Read this chapter as a roadmap: first understand the issues created by function approximation, then track how prediction, control, off-policy methods, eligibility traces, and policy-gradient methods receive different roles.
Three Warning Signs
The overview identifies three issues that appear when reinforcement learning uses function approximation: nonstationarity, bootstrapping, and delayed targets. At this stage, the important goal is not to memorize a detailed algorithm for each term. Instead, treat them as signals that the learning problem must be examined again rather than assumed to behave like conventional supervised learning.
The On-Policy Starting Point
The early sequence deliberately restricts attention to on-policy training. It begins with prediction using a given policy and then moves to control, whose aim is to find an approximation to the optimal policy. This ordering separates two learning objectives before the chapter sequence broadens to off-policy methods.
Classifying the First Three Chapters
A learner is reviewing Chapters 9, 10, and 11. Which learning setting or objective belongs to each chapter?
Chapter 9: Assign prediction because the policy is given and the value function is the part being approximated.
Chapter 10: Assign control because the aim is to find an approximation to the optimal policy.
Chapter 11: Assign off-policy methods because this chapter broadens the scope beyond the initial on-policy training sequence.
The sequence is prediction with a given policy, then control toward an optimal policy, then off-policy methods.
Prediction and Control
| Topic | What is given or emphasized | Learning objective |
|---|---|---|
| Prediction | A given policy | Approximate the value function |
| Control | A search for a better policy | Find an approximation to the optimal policy |
Prediction and control are related but not interchangeable. In prediction, the policy is already given, and the value function is the part being approximated. In control, the aim changes: the learner seeks an approximation to the optimal policy. This distinction explains why the early sequence treats prediction before control.
Mechanisms and Alternative Control
The later chapters do not all organize the material around the same kind of training setting. Chapter 12 focuses on eligibility traces, an algorithmic mechanism whose computational role is important in multistep reinforcement-learning methods. The overview states that eligibility traces can dramatically improve the computational properties of these methods in many cases. The final chapter presents policy-gradient methods as a different approach to control: they approximate the optimal policy directly.
Roadmap Mistakes
Treating function approximation as only a new representation
The overview says function approximation also introduces nonstationarity, bootstrapping, and delayed targets, including issues not normally encountered in conventional supervised learning.
Fix:
Use the three terms as warning signs that the learning objectives and methods need renewed analysis.Confusing prediction with control
Chapter 9 handles prediction with a given policy, while the move toward an optimal policy belongs to control.
Fix:
Associate prediction with approximating the value function for a given policy and control with seeking an approximation to the optimal policy.Treating eligibility traces as another broad training setting
Chapter 12 changes emphasis to an algorithmic mechanism and analyzes the computational role of eligibility traces in multistep methods.
Fix:
Classify eligibility traces as a mechanism that can improve the computational properties of multistep methods.Assuming policy-gradient methods must first form an approximate value function
The overview says policy-gradient methods need never form an approximate value function.
Fix:
Remember that policy-gradient methods directly approximate the optimal policy, while also noting that value-function approximation may improve efficiency.
Roadmap Check
Complete the roadmap without looking back: Chapter 9 addresses what kind of learning objective? What changes in Chapter 10? What broader setting appears in Chapter 11? What mechanism receives focused treatment in Chapter 12? What direct control approach appears in the final chapter?
Hints
- Start with the distinction between a given policy and a search for an optimal policy.
- Then separate training settings from algorithmic mechanisms.
- The final chapter uses a direct approach to approximating the optimal policy.
What do you think happens?
Which sequence best matches the chapter roadmap?
Reveal answer
Answer: Prediction with a given policy, control toward an optimal policy, off-policy methods, eligibility traces, then policy-gradient methods
The early sequence starts with on-policy prediction and control. Chapter 11 broadens to off-policy methods, Chapter 12 focuses on eligibility traces, and the final chapter presents policy-gradient methods as a direct control approach.
Chapter Takeaways
- The roadmap begins by warning that function approximation introduces nonstationarity, bootstrapping, and delayed targets. It then narrows attention to on-policy training: prediction approximates a value function for a given policy, while control seeks an approximation to the optimal policy. The sequence next broadens to off-policy methods, studies eligibility traces as a computational mechanism for multistep methods, and ends with policy-gradient methods as a direct approach to approximating the optimal policy.
Key Takeaways
- Function approximation introduces nonstationarity, bootstrapping, and delayed targets as issues requiring renewed analysis.
- The early sequence moves from on-policy prediction with a given policy to on-policy control toward an optimal policy.
- Chapter 11 broadens the scope to off-policy methods.
- Chapter 12 focuses on eligibility traces and their computational role in multistep methods.
- The final chapter presents policy-gradient methods as a direct approach to approximating the optimal policy.