Concepts / On-Policy and Off-Policy Learning

On-Policy and Off-Policy Learning

Exploration is necessary because the currently estimated best action may not be truly best.

  • Programming

When the Best Action Is Only an Estimate

An agent chooses actions using estimates of how valuable those actions are. The action that currently appears best may not actually be best. Learning therefore requires experience with alternatives, not just repeated use of the current favorite.

keeps selectingdoes not tryproducescannot produceStateAction Acurrently estimated bestReturns for AAction BalternativeReturns for Bno experience
What happens when an agent repeatedly selects the action it currently estimates as best?

What do you think happens?

Suppose Action A currently appears best and the agent stops trying Action B. What can the agent learn about B?

  • It can determine B's value from the returns it receives for A.
  • It obtains no returns for B, so B may never be recognized as better.
  • It immediately proves that A is truly best.
Reveal answer

Answer: It obtains no returns for B, so B may never be recognized as better.

Without trying an alternative action, the agent obtains no returns for that alternative. It therefore lacks the experience needed to compare the alternative with the action currently estimated as best.

Why Alternative Experience Matters

A common first thought in Monte Carlo control is to keep selecting an action once it appears to be best. This seems efficient because the agent avoids actions that currently look worse. The problem is that stopping exploration also stops the flow of returns for the alternatives. If Action B is never selected, the agent may never learn how B compares with Action A. The point is not that B is better; the point is that experience with B is required before the agent can determine that comparison.

A Missing Comparison

At a state, Action A is currently estimated as best. The agent repeatedly chooses A and never chooses Action B.

Current choice: The agent selects A because A currently has the higher estimate.

Missing experience: Because B is never selected, the agent obtains no returns for B.

Learning limitation: Without returns for B, the agent cannot determine from experience whether B is better than A.

Required response: The agent must retain some way to obtain experience with alternatives.

Exploration is part of the control problem rather than an optional activity added after learning is complete.

Exploring Starts and Coverage

Exploring starts address the lack of experience by attempting to begin episodes from different state-action pairs. This gives the learning process a way to obtain experience for pairs that ordinary action selection might not visit. Their purpose is broad coverage of state-action experience in simulated episodes.

begins an episodebegins an episodebegins an episodeState-action pair 1starting pairEpisode experiencebroader coverageState-action pair 2starting pairState-action pair 3starting pair
How do different starting states and actions help collect experience across otherwise unvisited state-action pairs?

When reasoning about exploring starts, ask which state-action pairs are receiving experience. The important question is not simply whether an episode begins, but whether starting episodes in different ways broadens the experience available to learning.

On-Policy Exploration

An on-policy method learns about the same policy that is choosing actions and continuing to explore. Exploration is therefore retained in the policy being learned. The method is not learning a policy that always chooses only the action currently estimated as best; it seeks the best policy under the requirement that the policy continues to explore.

selectsproducesinformscontinues learningExploring policypolicy being learnedActionchosen by policyExperiencereturnsUpdated policycontinues to explore
How does an on-policy method choose exploratory actions while learning about the policy making those choices?

The defining idea is alignment: the policy generating the experience is also the policy being evaluated and improved, and that policy continues to explore.

Off-Policy Exploration

An off-policy method separates the policy that generates experience from the policy being learned. The experience can come from an exploratory behavior policy, while a separate target policy is the policy being improved. This separation provides another way to preserve exploration in the data-generating process without making the behavior policy and the learned target policy the same policy.

generatesinformsBehavior policygenerates experienceExperiencestate-action returnsTarget policybeing learned
How does data move from an exploratory behavior policy to the separate target policy being learned?

Comparing the Two Approaches

QuestionOn-policy learningOff-policy learning
Which policy generates experience?The policy that continues to explore and is being learnedA behavior policy
Which policy is being evaluated or improved?The same exploring policyA separate target policy
How is exploration maintained?Exploration remains in the policy being learnedExploration is maintained by the behavior policy generating experience
one policyseparate rolesseparate rolesOn-policyExploring policyacts and is learnedBehavior policygenerates experienceOff-policyTarget policybeing learned
What is the difference between the policy generating experience and the policy being improved?

Simulation Versus Real Experience

Exploring starts are more practical in simulated episodes because a simulator can support deliberately beginning episodes from selected state-action pairs. In real experience, such starting conditions are unlikely: the system usually cannot conveniently arrange an episode to begin from every desired pair. Exploring starts therefore support broad coverage especially naturally in simulation, while real experience makes them difficult to rely on.

supportsmakes less likelySimulationselected startsState-action pairsbroad coverageSelected startsdifficult to arrangeReal experienceunlikely starts
Why can selected starting pairs be arranged more readily in simulation than in real experience?

Common Reasoning Mistakes

  • Assuming that the currently estimated best action is already known to be truly best.

    The estimate may be incorrect, and repeated selection of A prevents the agent from obtaining returns for alternatives such as B.

    Fix: Treat the current best action as an estimate that still requires comparison with alternatives.

  • Saying that Action B must be better simply because it needs to be explored.

    Exploration provides the experience needed for comparison; it does not establish the result of that comparison in advance.

    Fix: State only that experience with B is needed to determine how B compares with A.

  • Describing on-policy learning as learning a policy that stops exploring.

    On-policy learning retains exploration in the policy being learned.

    Fix: Remember that the policy being learned is also the policy that continues to explore.

  • Treating on-policy and off-policy learning as having the same policy roles.

    Off-policy learning separates the policy generating experience from the target policy being learned.

    Fix: Identify the behavior policy and target policy separately when describing off-policy learning.

  • Assuming exploring starts are equally practical in real experience.

    Exploring starts are unlikely in real experience, even though they can support broad coverage in simulated episodes.

    Fix: Connect exploring starts primarily with deliberately arranged simulated episodes.

Check Your Understanding

MEDIUM

Explain the difference between these two situations: first, a policy that explores while it is also the policy being learned; second, a behavior policy that generates exploratory experience for a separate target policy. Then explain why stopping exploration after one action appears best can prevent learning about alternatives.

Hints
  • Identify which policy generates the experience.
  • Identify which policy is being evaluated or improved.
  • Mention what happens to returns for an action that is never selected.
  1. A strong answer should say that on-policy learning evaluates and improves the same policy that continues to explore, whereas off-policy learning uses experience from a behavior policy to learn about a separate target policy. It should also say that an untried alternative produces no returns, so its value may never be recognized.

Key Takeaways

  • The currently estimated best action may not be truly best.
  • Selecting only that action prevents the agent from obtaining returns for alternatives.
  • Exploring starts attempt to provide experience across different state-action pairs and are especially practical in simulated episodes.
  • On-policy methods learn the policy that continues to explore.
  • Off-policy methods separate the exploratory behavior policy from the target policy being learned.

Key Takeaways

  • Exploration is necessary because the action currently estimated as best may not actually be best.
  • Without trying an alternative, the agent obtains no returns for it and may never learn its value.
  • Exploring starts broaden experience across state-action pairs, especially in simulated episodes.
  • On-policy learning retains exploration in the policy being learned.
  • Off-policy learning uses an exploratory behavior policy to generate experience for a separate target policy.