Concepts / On-policy Prediction

On-policy Prediction

Off-policy learning separates the policy that generates data from the policy being evaluated.

  • Programming

One State, Two Policies

Suppose you want to estimate the value of one Blackjack state, but the episodes available to you were generated by a policy different from the policy you want to evaluate. This is the central situation addressed by off-policy estimation. One policy supplies the experience, while another policy is the subject of the value estimate.

usesseparatesseparatesOn-policyone policyData policyevaluated policyOff-policytwo policiesBehavior policygenerates dataTarget policybeing evaluated
What changes when the policy generating experience is the same as the policy being evaluated versus when the two policies are different?

Off-policy learning separates the policy that generates data from the policy being evaluated. The data-generating policy is called the behavior policy, and the policy whose value is estimated is called the target policy.

The Blackjack Evaluation

The example evaluates one particular Blackjack state: the dealer is showing a deuce, the player has a sum of 13, and the player has a usable ace. Every sampled episode begins from this state, and the learning process uses complete episodes to estimate the value of that state.

choosesspecifiescompared with targetdetermines adjustmentRandom hit-or-stickgenerates episodesHit or stickchosen at randomStick on 20 or 21being evaluatedStickat 20 or 21Importance weightsadjust returns
Which policy generates each episode, which policy is evaluated, and how do their action choices determine the importance-sampling weights?
RolePolicy in the examplePurpose
Behavior policyRandomly chooses hit or stickGenerates the episodes
Target policySticks only when the player's sum is 20 or 21Is evaluated

The two policies have different roles, even though the same sampled episodes provide the data.

The important separation is therefore not between two Blackjack states. It is between two policies: the random policy supplies the experience, while the policy that sticks only on 20 or 21 is the policy whose value is estimated.

Reweighting Sampled Returns

An episode generated by the behavior policy is not automatically an unbiased description of what the target policy would produce, because the two policies make choices differently. Ordinary importance sampling addresses this difference by attaching an importance-sampling weight to the episode's return. The weight adjusts the contribution of that return so that the collection of behavior-policy episodes can estimate the target policy's value.

starts fromproducescompared with targetreweighted byadjustsBlackjack statedeuce, 13, usable aceSampled episodebehavior policyEpisode returnobserved outcomeTarget value estimateweighted returnImportance weightpolicy comparison
How does an importance-sampling ratio reweight a return generated by the behavior policy so it estimates the target policy's value?

One behavior-policy episode

An episode begins at the specified Blackjack state, is generated by the random hit-or-stick behavior policy, and produces a return. How does ordinary importance sampling use it?

1. Generate: The behavior policy generates a complete episode from the state with a dealer showing a deuce, player sum 13, and a usable ace.

2. Observe: The episode supplies a return, meaning the outcome used for the Monte Carlo estimate.

3. Compare: The actions selected while generating the episode are compared with the action choices specified by the target policy, which sticks only on 20 or 21.

4. Reweight: Ordinary importance sampling uses the resulting importance-sampling weight to adjust the episode's return before combining it with other weighted returns.

The episode remains behavior-policy data, but its weighted contribution is used to estimate the target policy's value.

Normalizing the Evidence

Weighted importance sampling uses the same idea of weighting returns, but it normalizes the weighted returns by the sum of their importance weights. In other words, it does not simply combine weighted returns directly; it also accounts for the total weight attached to the available evidence.

combinenormalize bycombinedivide byWeighted returnscombined directlyOrdinary estimateno normalizationWeighted returnssame evidenceSum of weightsnormalizing totalWeighted estimatenormalized
How are multiple weighted returns normalized by the sum of their importance weights, and how does this differ from ordinary importance sampling?
MethodHow returns are usedReported behavior
Ordinary importance samplingUses importance weights to combine returnsCan have high variance
Weighted importance samplingNormalizes weighted returns by the sum of the importance weightsTypically has lower initial error

Learning Across Episodes

samplesamplesamplereturnreturnreturnBlackjack statedeuce, 13, usable aceEpisode 1complete episodeState value estimatereturns combinedEpisode 2complete episodeEpisode Ncomplete episode
How do complete sampled episodes starting from the same Blackjack state contribute returns to that state's value estimate?

The experiment began every run with estimates of zero and then learned from 10,000 episodes. The process was repeated in 100 independent runs. At each episode count, the researchers averaged the squared error across those runs. This measures how far the learned estimate was from the value being estimated as learning progressed.

begins withapproachesbegins withapproachesOrdinary importancesamplingerror approaches zeroHigher initial errorreported patternWeighted importancesamplingerror approaches zeroLower initial errortypical patternError near zeroafter learning
How do the value estimates and errors of ordinary and weighted importance sampling change across successive episodes?

Both algorithms' error approached zero as more episodes were used. Weighted importance sampling had much lower error at the beginning. The source identifies this lower initial error as typical in practice, while ordinary importance sampling can have high variance. The result is not that ordinary importance sampling fails; rather, the two methods can behave differently early in learning even though both improve toward the desired value.

Mistakes in Policy Roles

  • Treating the random hit-or-stick policy as the policy being evaluated.

    In the example, the random policy is the behavior policy. The target policy is the policy that sticks only on 20 or 21.

    Fix: Ask two separate questions: Which policy generated the data? Which policy's value is being estimated?

  • Assuming that importance sampling changes the sampled episodes into target-policy episodes.

    The episode still came from the behavior policy. Importance sampling changes the episode's contribution to the estimate through a weight.

    Fix: Describe the data as behavior-policy data and the adjusted estimate as an estimate of the target policy's value.

  • Saying that weighted importance sampling ignores importance weights.

    Both methods use weighted returns. Weighted importance sampling additionally normalizes by the sum of the importance weights.

    Fix: Remember: ordinary importance sampling weights returns; weighted importance sampling weights and normalizes them.

  • Interpreting lower initial error as permanent superiority.

    The reported results say that both methods' error approached zero, while weighted importance sampling had much lower error at the beginning.

    Fix: Separate early learning behavior from the longer-run result reported for the experiment.

Check Your Understanding

MEDIUM

For the Blackjack experiment, explain in your own words why an episode generated by the random hit-or-stick policy can still contribute to an estimate of the policy that sticks only on 20 or 21. Then state the difference between ordinary and weighted importance sampling.

Hints
  • Start by naming the behavior policy and the target policy.
  • Explain what the importance-sampling weight does to an episode return.
  • Mention the sum of the importance weights when describing weighted importance sampling.

What do you think happens?

After many episodes in the reported experiment, what happened to the error of both ordinary and weighted importance sampling?

  • Both errors approached zero
  • Only ordinary importance sampling improved
  • Only weighted importance sampling improved
  • Both errors increased
Reveal answer

Answer: Both errors approached zero

The reported experiment found that both algorithms' error approached zero. Weighted importance sampling had much lower error at the beginning.

Key Takeaways

  1. Off-policy estimation uses data from a behavior policy to estimate the value of a different target policy.
  2. The Blackjack state in the example has a dealer showing a deuce, a player sum of 13, and a usable ace.
  3. The random hit-or-stick policy is the behavior policy, while sticking only on 20 or 21 is the target policy.
  4. Ordinary importance sampling weights episode returns; weighted importance sampling also normalizes those weighted returns by the sum of the importance weights.
  5. In the reported experiment, both methods' error approached zero, while weighted importance sampling typically had lower initial error.

Key Takeaways

  • Off-policy learning separates the policy that generates experience from the policy being evaluated.
  • In the Blackjack example, random hit-or-stick generates the data, and sticking only on 20 or 21 is evaluated.
  • Ordinary importance sampling reweights returns, whereas weighted importance sampling normalizes the weighted returns.
  • Both methods' error approached zero in the reported experiment, but weighted importance sampling had much lower initial error.