Concepts / Importance Sampling

Importance Sampling

Off-policy prediction separates the policy that generates experience from the policy whose value is estimated.

  • Programming

Two Policies, Two Roles

Imagine that an agent has already generated experience by following one policy, but you want to estimate the value of a different policy. Off-policy prediction addresses this situation. The policy that generated the episodes is the behavior policy. The policy whose value you want to estimate is the target policy.

The separation is important because the observed returns came from the behavior policy, not directly from the target policy. Importance sampling adjusts the influence of those returns so that the data can be used to estimate the target policy's value. The behavior policy determines what experience is available; the target policy determines how that experience is interpreted.

generatesexamine actionsweight returnsBehavior policygenerates episodesObserved episodesactions and returnsAction-probabilityratiosadjust influenceTarget policy valueestimated value function
How does experience generated by one policy flow into an estimate of another policy's value?

Weighting an Observed Action

For each action that actually occurred in an episode, importance sampling compares two probabilities: the probability that the target policy assigns to that action and the probability that the behavior policy assigns to it. The action's weight is based on the target-policy probability divided by the behavior-policy probability.

This comparison changes how strongly the observed return contributes to the estimate. If the target policy treats an observed action as relatively more likely than the behavior policy does, the return receives more influence. If the target policy treats it as relatively less likely, the return receives less influence. The ratio therefore adjusts behavior-policy experience toward the target policy being evaluated.

evaluate under targetevaluate under behaviornumeratordenominatorweightsObserved actionfrom an episodeTarget probabilityprobability of actionBehavior probabilityprobability of actionProbability ratiotarget divided by behaviorWeighted returnadjusted influence
How is each observed action's probability under the target policy compared with its probability under the behavior policy, and how does that ratio change the return's weight?

Accumulating a Trajectory Weight

A Two-Action Episode

An observed episode contains two actions. For the first action, the target policy assigns probability 0.8 and the behavior policy assigns probability 0.4. For the second action, the target policy assigns probability 0.5 and the behavior policy assigns probability 0.5. Determine how the episode's return is weighted.

Compare the first action: The target policy assigns twice as much probability to the first observed action as the behavior policy does, so the first action contributes a ratio of 2 to the episode's overall weight.

Compare the second action: Both policies assign the same probability to the second observed action, so this action contributes a ratio of 1.

Accumulate the trajectory weight: The action-probability ratios are accumulated across the sequence of decisions. Here, the two ratios are combined multiplicatively, giving an overall trajectory weight of 2.

Apply the weight: The observed return is given a weight of 2 when it contributes to the estimate of the target policy's value.

The episode's return receives an overall weight of 2. The example illustrates that the weight reflects the entire sequence of observed actions, not just one action.

next decisionaccumulatecombineweightFirst actionratio 2Second actionratio 1Episode weightcombined weight 2Observed returnweighted contribution
How do action-probability ratios accumulate across a sequence of decisions to produce a single weight for a sampled return?

The example uses a short episode to make the mechanism visible. In general, the learning process examines the actions that occurred in each episode, compares their probabilities under the two policies, uses the comparisons to weight returns, and combines the weighted information into an estimate of the target policy's value function.

Two Ways to Average

EstimatorHow returns are combinedMain propertyPractical implication
Ordinary importance samplingTakes a simple average of the weighted returnsUnbiased, but variance can be larger and possibly infiniteThe estimate can be highly variable
Weighted importance samplingUses a weighted average, normalizing by the sum of the weightsHas finite varianceGenerally preferred in practice

Ordinary importance sampling averages the weighted returns directly. Its key advantage is unbiasedness: in the description provided here, it produces unbiased estimates. Its drawback is variance. That variance can be larger and can possibly be infinite, so individual estimates may vary substantially.

Weighted importance sampling instead forms a weighted average. The weights affect both the returns' contributions and the normalization of the average. This estimator has finite variance and is preferred in practice, making it the usual practical choice when selecting between the two forms described here.

propertypropertyOrdinary samplingsimple averageWeighted samplingweighted averageUnbiased estimatevariance may be infiniteFinite variancepreferred in practice
What is the difference between averaging importance-weighted returns directly and normalizing them by the sum of their weights?

Why Practice Favors Weighted Estimates

The practical choice follows from the estimators' variance properties. Ordinary importance sampling is unbiased, but its variance may be large or even infinite. Weighted importance sampling has finite variance, so its estimates are more suitable for practical value estimation according to the source material.

can producehasDirect weightedaverageordinary samplingNormalized weightedaverageweighted samplingLarge variancepossibly infiniteFinite variancepreferred in practice
How do normalization and extreme trajectory weights affect estimator stability, variance, and practical value estimates?

Common Reasoning Errors

  • Treating the behavior policy as the policy being evaluated

    Off-policy prediction separates the policy that supplies experience from the policy whose value is estimated.

    Fix: Identify the behavior policy as the data source and the target policy as the policy being evaluated.

  • Ignoring the behavior-policy probability

    Importance sampling adjusts for the policy difference by comparing the target-policy probability with the behavior-policy probability.

    Fix: Use the target probability relative to the behavior probability for every observed action.

  • Applying a ratio to only one action in a multi-action episode

    The action-probability comparisons accumulate across the sequence of decisions in the episode.

    Fix: Account for the observed actions across the trajectory before applying the resulting weight to the return.

  • Claiming that ordinary importance sampling is always the practical choice because it is unbiased

    Ordinary importance sampling can have larger, possibly infinite, variance.

    Fix: Recognize that weighted importance sampling has finite variance and is preferred in practice.

Check Your Understanding

MEDIUM

An episode was generated by a behavior policy, but the learning goal is to estimate the value of a target policy. Explain the role of each policy, describe what is compared for every observed action, and choose between ordinary and weighted importance sampling for practical use. Justify the choice using the variance properties of the two estimators.

Hints
  • The behavior policy supplies the observed experience.
  • The target policy is the policy whose value is being estimated.
  • Compare the action probabilities assigned by the two policies.
  • Weighted importance sampling is preferred because it has finite variance.

What do you think happens?

Suppose an observed action is assigned a higher probability by the target policy than by the behavior policy. Will that action make the observed return more influential or less influential in the target-policy estimate?

  • More influential
  • Less influential
  • It has no effect
Reveal answer

Answer: More influential

The target-policy probability is larger relative to the behavior-policy probability, so the action-probability ratio gives the return more weight.

Key Takeaways

  1. Off-policy prediction estimates a target policy's value using episodes generated by a behavior policy.
  2. Importance sampling compares target-policy and behavior-policy probabilities for the actions that actually occurred.
  3. The action-probability comparisons accumulate across an episode and determine how strongly its return contributes.
  4. Ordinary importance sampling averages weighted returns directly and is unbiased, but its variance can be larger or possibly infinite.
  5. Weighted importance sampling uses a weighted average, has finite variance, and is preferred in practice.

Key Takeaways

  • Behavior and target policies have different roles: one generates experience, while the other is evaluated.
  • Importance sampling adjusts behavior-policy returns using action-probability ratios.
  • Ratios from the actions in a trajectory combine to determine the trajectory's return weight.
  • Ordinary importance sampling is unbiased but may have large or infinite variance.
  • Weighted importance sampling has finite variance and is generally preferred in practice.