On-Policy and Off-Policy Learning
Exploration is necessary because the currently estimated best action may not be truly best.
When the Best Action Is Only an Estimate
An agent chooses actions using estimates of how valuable those actions are. The action that currently appears best may not actually be best. Learning therefore requires experience with alternatives, not just repeated use of the current favorite.
What do you think happens?
Suppose Action A currently appears best and the agent stops trying Action B. What can the agent learn about B?
Reveal answer
Answer: It obtains no returns for B, so B may never be recognized as better.
Without trying an alternative action, the agent obtains no returns for that alternative. It therefore lacks the experience needed to compare the alternative with the action currently estimated as best.
Why Alternative Experience Matters
A common first thought in Monte Carlo control is to keep selecting an action once it appears to be best. This seems efficient because the agent avoids actions that currently look worse. The problem is that stopping exploration also stops the flow of returns for the alternatives. If Action B is never selected, the agent may never learn how B compares with Action A. The point is not that B is better; the point is that experience with B is required before the agent can determine that comparison.
A Missing Comparison
At a state, Action A is currently estimated as best. The agent repeatedly chooses A and never chooses Action B.
Current choice: The agent selects A because A currently has the higher estimate.
Missing experience: Because B is never selected, the agent obtains no returns for B.
Learning limitation: Without returns for B, the agent cannot determine from experience whether B is better than A.
Required response: The agent must retain some way to obtain experience with alternatives.
Exploration is part of the control problem rather than an optional activity added after learning is complete.
Exploring Starts and Coverage
Exploring starts address the lack of experience by attempting to begin episodes from different state-action pairs. This gives the learning process a way to obtain experience for pairs that ordinary action selection might not visit. Their purpose is broad coverage of state-action experience in simulated episodes.
When reasoning about exploring starts, ask which state-action pairs are receiving experience. The important question is not simply whether an episode begins, but whether starting episodes in different ways broadens the experience available to learning.
On-Policy Exploration
An on-policy method learns about the same policy that is choosing actions and continuing to explore. Exploration is therefore retained in the policy being learned. The method is not learning a policy that always chooses only the action currently estimated as best; it seeks the best policy under the requirement that the policy continues to explore.
The defining idea is alignment: the policy generating the experience is also the policy being evaluated and improved, and that policy continues to explore.
Off-Policy Exploration
An off-policy method separates the policy that generates experience from the policy being learned. The experience can come from an exploratory behavior policy, while a separate target policy is the policy being improved. This separation provides another way to preserve exploration in the data-generating process without making the behavior policy and the learned target policy the same policy.
Comparing the Two Approaches
| Question | On-policy learning | Off-policy learning |
|---|---|---|
| Which policy generates experience? | The policy that continues to explore and is being learned | A behavior policy |
| Which policy is being evaluated or improved? | The same exploring policy | A separate target policy |
| How is exploration maintained? | Exploration remains in the policy being learned | Exploration is maintained by the behavior policy generating experience |
Simulation Versus Real Experience
Exploring starts are more practical in simulated episodes because a simulator can support deliberately beginning episodes from selected state-action pairs. In real experience, such starting conditions are unlikely: the system usually cannot conveniently arrange an episode to begin from every desired pair. Exploring starts therefore support broad coverage especially naturally in simulation, while real experience makes them difficult to rely on.
Common Reasoning Mistakes
Assuming that the currently estimated best action is already known to be truly best.
The estimate may be incorrect, and repeated selection of A prevents the agent from obtaining returns for alternatives such as B.
Fix:
Treat the current best action as an estimate that still requires comparison with alternatives.Saying that Action B must be better simply because it needs to be explored.
Exploration provides the experience needed for comparison; it does not establish the result of that comparison in advance.
Fix:
State only that experience with B is needed to determine how B compares with A.Describing on-policy learning as learning a policy that stops exploring.
On-policy learning retains exploration in the policy being learned.
Fix:
Remember that the policy being learned is also the policy that continues to explore.Treating on-policy and off-policy learning as having the same policy roles.
Off-policy learning separates the policy generating experience from the target policy being learned.
Fix:
Identify the behavior policy and target policy separately when describing off-policy learning.Assuming exploring starts are equally practical in real experience.
Exploring starts are unlikely in real experience, even though they can support broad coverage in simulated episodes.
Fix:
Connect exploring starts primarily with deliberately arranged simulated episodes.
Check Your Understanding
Explain the difference between these two situations: first, a policy that explores while it is also the policy being learned; second, a behavior policy that generates exploratory experience for a separate target policy. Then explain why stopping exploration after one action appears best can prevent learning about alternatives.
Hints
- Identify which policy generates the experience.
- Identify which policy is being evaluated or improved.
- Mention what happens to returns for an action that is never selected.
- A strong answer should say that on-policy learning evaluates and improves the same policy that continues to explore, whereas off-policy learning uses experience from a behavior policy to learn about a separate target policy. It should also say that an untried alternative produces no returns, so its value may never be recognized.
Key Takeaways
- The currently estimated best action may not be truly best.
- Selecting only that action prevents the agent from obtaining returns for alternatives.
- Exploring starts attempt to provide experience across different state-action pairs and are especially practical in simulated episodes.
- On-policy methods learn the policy that continues to explore.
- Off-policy methods separate the exploratory behavior policy from the target policy being learned.
Key Takeaways
- Exploration is necessary because the action currently estimated as best may not actually be best.
- Without trying an alternative, the agent obtains no returns for it and may never learn its value.
- Exploring starts broaden experience across state-action pairs, especially in simulated episodes.
- On-policy learning retains exploration in the policy being learned.
- Off-policy learning uses an exploratory behavior policy to generate experience for a separate target policy.