Importance Sampling
Off-policy prediction separates the policy that generates experience from the policy whose value is estimated.
Two Policies, Two Roles
Imagine that an agent has already generated experience by following one policy, but you want to estimate the value of a different policy. Off-policy prediction addresses this situation. The policy that generated the episodes is the behavior policy. The policy whose value you want to estimate is the target policy.
The separation is important because the observed returns came from the behavior policy, not directly from the target policy. Importance sampling adjusts the influence of those returns so that the data can be used to estimate the target policy's value. The behavior policy determines what experience is available; the target policy determines how that experience is interpreted.
Weighting an Observed Action
For each action that actually occurred in an episode, importance sampling compares two probabilities: the probability that the target policy assigns to that action and the probability that the behavior policy assigns to it. The action's weight is based on the target-policy probability divided by the behavior-policy probability.
This comparison changes how strongly the observed return contributes to the estimate. If the target policy treats an observed action as relatively more likely than the behavior policy does, the return receives more influence. If the target policy treats it as relatively less likely, the return receives less influence. The ratio therefore adjusts behavior-policy experience toward the target policy being evaluated.
Accumulating a Trajectory Weight
A Two-Action Episode
An observed episode contains two actions. For the first action, the target policy assigns probability 0.8 and the behavior policy assigns probability 0.4. For the second action, the target policy assigns probability 0.5 and the behavior policy assigns probability 0.5. Determine how the episode's return is weighted.
Compare the first action: The target policy assigns twice as much probability to the first observed action as the behavior policy does, so the first action contributes a ratio of 2 to the episode's overall weight.
Compare the second action: Both policies assign the same probability to the second observed action, so this action contributes a ratio of 1.
Accumulate the trajectory weight: The action-probability ratios are accumulated across the sequence of decisions. Here, the two ratios are combined multiplicatively, giving an overall trajectory weight of 2.
Apply the weight: The observed return is given a weight of 2 when it contributes to the estimate of the target policy's value.
The episode's return receives an overall weight of 2. The example illustrates that the weight reflects the entire sequence of observed actions, not just one action.
The example uses a short episode to make the mechanism visible. In general, the learning process examines the actions that occurred in each episode, compares their probabilities under the two policies, uses the comparisons to weight returns, and combines the weighted information into an estimate of the target policy's value function.
Two Ways to Average
| Estimator | How returns are combined | Main property | Practical implication |
|---|---|---|---|
| Ordinary importance sampling | Takes a simple average of the weighted returns | Unbiased, but variance can be larger and possibly infinite | The estimate can be highly variable |
| Weighted importance sampling | Uses a weighted average, normalizing by the sum of the weights | Has finite variance | Generally preferred in practice |
Ordinary importance sampling averages the weighted returns directly. Its key advantage is unbiasedness: in the description provided here, it produces unbiased estimates. Its drawback is variance. That variance can be larger and can possibly be infinite, so individual estimates may vary substantially.
Weighted importance sampling instead forms a weighted average. The weights affect both the returns' contributions and the normalization of the average. This estimator has finite variance and is preferred in practice, making it the usual practical choice when selecting between the two forms described here.
Why Practice Favors Weighted Estimates
The practical choice follows from the estimators' variance properties. Ordinary importance sampling is unbiased, but its variance may be large or even infinite. Weighted importance sampling has finite variance, so its estimates are more suitable for practical value estimation according to the source material.
Common Reasoning Errors
Treating the behavior policy as the policy being evaluated
Off-policy prediction separates the policy that supplies experience from the policy whose value is estimated.
Fix:
Identify the behavior policy as the data source and the target policy as the policy being evaluated.Ignoring the behavior-policy probability
Importance sampling adjusts for the policy difference by comparing the target-policy probability with the behavior-policy probability.
Fix:
Use the target probability relative to the behavior probability for every observed action.Applying a ratio to only one action in a multi-action episode
The action-probability comparisons accumulate across the sequence of decisions in the episode.
Fix:
Account for the observed actions across the trajectory before applying the resulting weight to the return.Claiming that ordinary importance sampling is always the practical choice because it is unbiased
Ordinary importance sampling can have larger, possibly infinite, variance.
Fix:
Recognize that weighted importance sampling has finite variance and is preferred in practice.
Check Your Understanding
An episode was generated by a behavior policy, but the learning goal is to estimate the value of a target policy. Explain the role of each policy, describe what is compared for every observed action, and choose between ordinary and weighted importance sampling for practical use. Justify the choice using the variance properties of the two estimators.
Hints
- The behavior policy supplies the observed experience.
- The target policy is the policy whose value is being estimated.
- Compare the action probabilities assigned by the two policies.
- Weighted importance sampling is preferred because it has finite variance.
What do you think happens?
Suppose an observed action is assigned a higher probability by the target policy than by the behavior policy. Will that action make the observed return more influential or less influential in the target-policy estimate?
Reveal answer
Answer: More influential
The target-policy probability is larger relative to the behavior-policy probability, so the action-probability ratio gives the return more weight.
Key Takeaways
- Off-policy prediction estimates a target policy's value using episodes generated by a behavior policy.
- Importance sampling compares target-policy and behavior-policy probabilities for the actions that actually occurred.
- The action-probability comparisons accumulate across an episode and determine how strongly its return contributes.
- Ordinary importance sampling averages weighted returns directly and is unbiased, but its variance can be larger or possibly infinite.
- Weighted importance sampling uses a weighted average, has finite variance, and is preferred in practice.
Key Takeaways
- Behavior and target policies have different roles: one generates experience, while the other is evaluated.
- Importance sampling adjusts behavior-policy returns using action-probability ratios.
- Ratios from the actions in a trajectory combine to determine the trajectory's return weight.
- Ordinary importance sampling is unbiased but may have large or infinite variance.
- Weighted importance sampling has finite variance and is generally preferred in practice.