Concepts / Ordinary Importance Sampling

Ordinary Importance Sampling

The behavior policy and target policy may generate and evaluate different action choices.

  • Programming

Two Policies, Two Jobs

Off-policy learning separates episode generation from value evaluation. A behavior policy generates the episodes, while a target policy is the policy whose value we want to estimate. Because these policies may choose actions differently, importance sampling tries to correct the observed returns according to how compatible each episode is with the target policy.

The central danger is that the correction can give a few episodes extremely large influence. In the looping example, ordinary importance sampling can therefore have infinite variance even though the target-policy value is 1.

Building an Episode Ratio

Importance sampling considers the action choices made along an entire episode. At each step, the behavior policy's action probability is compared with the target policy's action probability. The episode correction combines these per-step comparisons across the trajectory. Thus, a mismatch at one step affects the episode's correction, and repeated mismatches across a long episode can produce a very different overall weight from that of a short episode.

observed choiceevaluated choicecombine across stepsBehavior policyaction probabilityPer-step comparisonone actionEpisode ratiocombined correctionTarget policyaction probability
How do the behavior policy's and target policy's action probabilities combine across an episode to produce an importance-sampling ratio?

What Looping Does to the Weight

Episodes that return to s

Follow the looping episode pattern used in the example: an episode takes some number of back transitions that return to s, followed by a back transition that terminates.

Episodes ending with end: These episodes have an importance-sampling ratio of zero because the target policy never chooses end. Their scaled contribution is therefore zero.

A short looping episode: An episode with a particular number of back transitions receives the correction associated with that length. Its return is 1.

A longer looping episode: An episode with more back transitions receives a different correction. The return is still 1, but the longer trajectory changes the importance-sampling correction.

All possible lengths: To examine the variance, consider every possible looping length, multiply each length's occurrence probability by the square of its importance-sampling ratio, and add the results. In this example, the total is infinite.

The target-policy value is 1, but ordinary importance sampling can remain unstable because the squared scaled return does not have a finite expectation.

loop againend through backapply episode correctionOne backreturn to sMore backslonger episodeTerminating backreturn 1Episode correctionlength-dependent
What happens to the importance-sampling weight as an episode repeatedly loops and combines more action-probability corrections?

The important contrast is between return and influence. Every contributing looping episode in this example has return 1, but different episode lengths receive different corrections. Rare long trajectories can therefore dominate the spread of ordinary estimates. The issue is not that the target value changes; it is that the scaled returns can vary so severely that their variance is infinite.

Ordinary and Weighted Estimators

Ordinary importance samplingWeighted importance sampling
Uses the scaled returns directly in the estimate.Normalizes the weighted returns into a weighted average.
Can remain unstable when rare episodes receive very large corrections.Ignores zero-weight inconsistent returns and averages the remaining returns with their weights.
Can have infinite variance in the looping example.In the example, becomes exactly 1 once a contributing episode ending through back appears.
average raw influencenormalize nonzero weightsOrdinary samplingscaled returnsUnstable estimateinfinite varianceWeighted samplingnormalized returnsEstimate 1after a back-ending episode
What is the difference between averaging raw weighted returns and normalizing the weights in the looping example?

Weighted importance sampling changes the normalization rather than treating every scaled return as a direct contribution to an ordinary average. Episodes ending with end have zero weight and are ignored. Every remaining contributing return in this particular example is 1. Therefore, once the data contains an episode that ends through back, the weighted average is exactly 1 and remains there in the example.

Mistakes in Reading the Example

  • Assuming that a return of 1 means every episode has the same influence.

    Importance sampling scales returns according to the compatibility of the episode with the target policy, so equal returns can still have different scaled contributions.

    Fix: Track both the return and the episode's correction.

  • Treating episodes that end with end as ordinary evidence for the target policy.

    Its scaled contribution is zero.

    Fix: Recognize that these episodes contribute nothing to the scaled quantity and receive zero weight in weighted importance sampling.

  • Concluding that ordinary importance sampling is stable because its target value is 1.

    The value and the estimator's spread are different properties.

    Fix: Examine the squared scaled returns when reasoning about variance.

  • Saying that weighted importance sampling changes the return of every episode.

    The stabilization comes from the weighted average and the zero-weight treatment of inconsistent returns.

    Fix: Describe the change as normalization of weighted returns.

Check Your Reasoning

MEDIUM

In the looping example, explain why an episode ending with end contributes zero, while an episode containing back transitions and ending through back can contribute to the estimate. Then explain why a longer looping episode can increase instability even though its return is still 1.

Hints
  • Start with the target policy's action choice for end.
  • Separate the return from the importance-sampling correction.
  • Connect the different episode lengths to the squared scaled-return calculation.

What do you think happens?

Once the data contains an episode that ends through back, what does weighted importance sampling produce in this particular example?

  • It remains undefined because the looping episodes have different lengths.
  • It becomes exactly 1 and remains there in the example.
  • It gives zero because episodes ending with end have zero weight.
Reveal answer

Answer: It becomes exactly 1 and remains there in the example.

Episodes ending with end have zero weight, and every remaining contributing return is 1. The weighted average is therefore exactly 1 once a back-ending episode appears.

Main Takeaways

  1. Off-policy learning uses one policy to generate episodes and another policy whose value is being estimated.
  2. Importance sampling combines action-probability comparisons across an episode, so looping trajectories can receive very different corrections from short trajectories.
  3. In the looping example, the target value is 1, but ordinary importance sampling has infinite variance because the squared scaled returns have an infinite total expectation.
  4. Weighted importance sampling normalizes the weighted returns and ignores zero-weight inconsistent episodes.
  5. In this example, every nonzero-weight contributing return is 1, so the weighted estimate becomes exactly 1 after a back-ending episode appears.

Key Takeaways

  • Ordinary importance sampling corrects for the mismatch between a behavior policy and a target policy by scaling episode returns.
  • Repeated loops create longer trajectories with different episode corrections, producing unusually large variation in scaled returns.
  • The looping example has target value 1 but infinite variance for the ordinary estimator.
  • Weighted importance sampling normalizes the weights, ignores zero-weight inconsistent returns, and becomes exactly 1 in this example once a contributing episode appears.