Weighted Importance Sampling
The behavior policy and target policy may generate and evaluate different action choices.
When Matching Policies Matters
Off-policy learning separates two jobs. A behavior policy generates episodes, while a target policy is the policy whose value we want to estimate. Because the two policies may choose different actions, an observed return may not represent the target policy directly. Importance sampling corrects for this mismatch by assigning each return an importance-sampling weight based on how compatible its episode is with the target policy.
How Loops Magnify Corrections
Consider the looping example. Some episodes end with end. Their importance-sampling ratio is zero because the target policy never chooses end, so they contribute nothing to the scaled quantity. Other episodes contain some number of back transitions that return to s and then finish with a back transition that terminates. Each of these episodes has return 1, but episodes with different loop lengths receive different importance-sampling corrections.
The important effect is multiplication across the episode. Longer looping trajectories accumulate a different correction from shorter trajectories. To determine the variance of the ordinary estimator, the analysis considers every possible loop length, multiplies each episode's occurrence probability by the square of its ratio, and adds the results. In this example, that total is infinite. Thus, ordinary importance sampling can have infinite variance even though the target-policy value is 1.
Two Ways to Combine Returns
Ordinary importance sampling forms a simple average of the scaled returns. It is unbiased: across repeated sampling, its expected estimate is the target value. However, that guarantee does not limit the size of fluctuations in individual estimates. If importance-sampling ratios can become very large, the variance can be unbounded.
Weighted importance sampling instead divides the total of the weighted returns by the total of the weights. This is a normalized weighted average. Episodes with zero weight are ignored, while episodes with nonzero weight share the estimate according to their relative weights.
Why Normalization Stabilizes the Example
The looping example
Compare the two estimators when episodes ending with end have zero weight and every remaining contributing episode has return 1.
Identify zero-weight episodes: Episodes ending with end are inconsistent with the target policy because the target policy never chooses end. Their weight is zero, so weighted importance sampling leaves them out of its weighted average.
Identify contributing episodes: Episodes containing back transitions and ending through back have nonzero weight in the example. Every one of these contributing returns is 1, even though different loop lengths receive different corrections.
Apply weighted normalization: The weighted average combines only the nonzero-weight returns. Since all of those returns equal 1, the weighted estimate is exactly 1 as soon as at least one such episode has been observed.
Compare ordinary averaging: Ordinary importance sampling averages scaled returns without the same normalization. Different loop lengths can therefore give some observations unusually large influence, producing unstable estimates even though the target-policy value is 1.
In this example, weighted importance sampling becomes exactly 1 after the first nonzero-weight episode and remains there, while ordinary importance sampling can remain unstable.
The One-Return Boundary Case
With one observed return and a nonzero importance-sampling ratio, weighted importance sampling returns the same value as the observed return because the ratio appears in both the numerator and the denominator and cancels. This does not mean that every one-return case is defined by cancellation. If the total weight is zero, the weighted estimate is defined to be zero.
Assuming unbiased means stable.
Unbiasedness describes the expected estimate across repeated sampling, not the size of fluctuations in an individual estimate.
Fix:
Compare both bias and variance when evaluating an estimator.Treating weighted importance sampling as ordinary averaging with different notation.
Normalization changes the estimator's statistical behavior: it becomes biased but has bounded variance.
Fix:
Remember that the two estimators make different bias and variance trade-offs.Treating a zero-weight episode as a useful target-policy return.
The episode is ignored by the weighted calculation because it is inconsistent with the target policy.
Fix:
Keep the observed return separate from the weight assigned to that return.Applying one-return cancellation when the total weight is zero.
Cancellation requires a nonzero ratio and a nonzero denominator.
Fix:
Check whether the total weight is zero before applying the single-return reasoning.
Choosing an Estimator
Weighted importance sampling is usually preferred in practice because it generally has dramatically lower variance. No single return receives a normalized weight larger than one, and when returns are bounded, the variance converges to zero even when the importance-sampling ratios themselves have infinite variance. The price is bias, although that bias approaches zero as more data is collected.
Ordinary importance sampling still has a useful role. Its simpler form is easier to extend to approximate methods that use function approximation. Therefore, the choice is not that ordinary importance sampling is always incorrect. It is that weighted importance sampling is often the more practical choice when estimate stability matters most.
An episode has a nonzero importance-sampling weight, and it is the only observed episode so far. Explain why weighted importance sampling returns the episode's observed return rather than a differently scaled value. Then explain what changes if every observed episode has zero weight.
Hints
- Track the same nonzero weight in both the weighted-return total and the total-weight denominator.
- For the second case, use the estimator's explicit zero-denominator rule.
- Ordinary importance sampling uses a simple average and is unbiased, but looping episodes can make its variance infinite.
- Repeated action choices create different episode-level corrections, so a few long trajectories can have unusually large influence.
- Weighted importance sampling divides total weighted returns by total weight and ignores zero-weight inconsistent returns.
- Weighted importance sampling is biased with bias converging to zero, but it has bounded variance and is usually more stable.
- Ordinary importance sampling remains useful because its simpler form is easier to extend to approximate methods using function approximation.
Key Takeaways
- Ordinary importance sampling is unbiased but can have infinite variance when looping episodes produce very large corrections.
- Weighted importance sampling normalizes weighted returns by the total weight.
- In the looping example, every contributing weighted return is 1, so the weighted estimate becomes exactly 1 after a nonzero-weight episode is observed.
- Weighted importance sampling trades bias for bounded variance, with bias converging to zero as more data is collected.
- Weighted importance sampling is usually preferred for stability, while ordinary importance sampling can remain useful for simpler approximate extensions.