Ordinary Importance Sampling
The behavior policy and target policy may generate and evaluate different action choices.
Two Policies, Two Jobs
Off-policy learning separates episode generation from value evaluation. A behavior policy generates the episodes, while a target policy is the policy whose value we want to estimate. Because these policies may choose actions differently, importance sampling tries to correct the observed returns according to how compatible each episode is with the target policy.
The central danger is that the correction can give a few episodes extremely large influence. In the looping example, ordinary importance sampling can therefore have infinite variance even though the target-policy value is 1.
Building an Episode Ratio
Importance sampling considers the action choices made along an entire episode. At each step, the behavior policy's action probability is compared with the target policy's action probability. The episode correction combines these per-step comparisons across the trajectory. Thus, a mismatch at one step affects the episode's correction, and repeated mismatches across a long episode can produce a very different overall weight from that of a short episode.
What Looping Does to the Weight
Episodes that return to s
Follow the looping episode pattern used in the example: an episode takes some number of back transitions that return to s, followed by a back transition that terminates.
Episodes ending with end: These episodes have an importance-sampling ratio of zero because the target policy never chooses end. Their scaled contribution is therefore zero.
A short looping episode: An episode with a particular number of back transitions receives the correction associated with that length. Its return is 1.
A longer looping episode: An episode with more back transitions receives a different correction. The return is still 1, but the longer trajectory changes the importance-sampling correction.
All possible lengths: To examine the variance, consider every possible looping length, multiply each length's occurrence probability by the square of its importance-sampling ratio, and add the results. In this example, the total is infinite.
The target-policy value is 1, but ordinary importance sampling can remain unstable because the squared scaled return does not have a finite expectation.
The important contrast is between return and influence. Every contributing looping episode in this example has return 1, but different episode lengths receive different corrections. Rare long trajectories can therefore dominate the spread of ordinary estimates. The issue is not that the target value changes; it is that the scaled returns can vary so severely that their variance is infinite.
Ordinary and Weighted Estimators
| Ordinary importance sampling | Weighted importance sampling |
|---|---|
| Uses the scaled returns directly in the estimate. | Normalizes the weighted returns into a weighted average. |
| Can remain unstable when rare episodes receive very large corrections. | Ignores zero-weight inconsistent returns and averages the remaining returns with their weights. |
| Can have infinite variance in the looping example. | In the example, becomes exactly 1 once a contributing episode ending through back appears. |
Weighted importance sampling changes the normalization rather than treating every scaled return as a direct contribution to an ordinary average. Episodes ending with end have zero weight and are ignored. Every remaining contributing return in this particular example is 1. Therefore, once the data contains an episode that ends through back, the weighted average is exactly 1 and remains there in the example.
Mistakes in Reading the Example
Assuming that a return of 1 means every episode has the same influence.
Importance sampling scales returns according to the compatibility of the episode with the target policy, so equal returns can still have different scaled contributions.
Fix:
Track both the return and the episode's correction.Treating episodes that end with end as ordinary evidence for the target policy.
Its scaled contribution is zero.
Fix:
Recognize that these episodes contribute nothing to the scaled quantity and receive zero weight in weighted importance sampling.Concluding that ordinary importance sampling is stable because its target value is 1.
The value and the estimator's spread are different properties.
Fix:
Examine the squared scaled returns when reasoning about variance.Saying that weighted importance sampling changes the return of every episode.
The stabilization comes from the weighted average and the zero-weight treatment of inconsistent returns.
Fix:
Describe the change as normalization of weighted returns.
Check Your Reasoning
In the looping example, explain why an episode ending with end contributes zero, while an episode containing back transitions and ending through back can contribute to the estimate. Then explain why a longer looping episode can increase instability even though its return is still 1.
Hints
- Start with the target policy's action choice for end.
- Separate the return from the importance-sampling correction.
- Connect the different episode lengths to the squared scaled-return calculation.
What do you think happens?
Once the data contains an episode that ends through back, what does weighted importance sampling produce in this particular example?
Reveal answer
Answer: It becomes exactly 1 and remains there in the example.
Episodes ending with end have zero weight, and every remaining contributing return is 1. The weighted average is therefore exactly 1 once a back-ending episode appears.
Main Takeaways
- Off-policy learning uses one policy to generate episodes and another policy whose value is being estimated.
- Importance sampling combines action-probability comparisons across an episode, so looping trajectories can receive very different corrections from short trajectories.
- In the looping example, the target value is 1, but ordinary importance sampling has infinite variance because the squared scaled returns have an infinite total expectation.
- Weighted importance sampling normalizes the weighted returns and ignores zero-weight inconsistent episodes.
- In this example, every nonzero-weight contributing return is 1, so the weighted estimate becomes exactly 1 after a back-ending episode appears.
Key Takeaways
- Ordinary importance sampling corrects for the mismatch between a behavior policy and a target policy by scaling episode returns.
- Repeated loops create longer trajectories with different episode corrections, producing unusually large variation in scaled returns.
- The looping example has target value 1 but infinite variance for the ordinary estimator.
- Weighted importance sampling normalizes the weights, ignores zero-weight inconsistent returns, and becomes exactly 1 in this example once a contributing episode appears.