Concepts / Off-policy Prediction with Importance Sampling

Off-policy Prediction with Importance Sampling

Importance sampling lets off-policy methods use behavior-policy samples to estimate quantities associated with a target policy.

  • Programming

One Experience, Two Policies

Off-policy prediction begins with a mismatch. Experience is collected under one policy, called the behavior policy, but the quantity we want to estimate belongs to another policy, called the target policy. Importance sampling makes those behavior-policy experiences useful by adjusting their influence according to how likely the same trajectories would have been under the target policy.

The central question is not whether a trajectory was observed. It is how likely that same trajectory would be under the target policy compared with the behavior policy.

generatescompare probabilitiesadjust influencecontributes toBehavior policyObserved trajectoryImportance-samplingratioWeighted returnTarget-policyestimate
How can data generated by the behavior policy be reweighted so that it estimates the return associated with the target policy?

Building a Trajectory Probability

Start at a particular state S_t. A subsequent trajectory includes the action selected at that state, the next state, the next action, and continuing state-action pairs until the endpoint. The probability of this complete trajectory contains factors from two sources: the policy's probabilities for selecting the observed actions and the MDP's transition probabilities for reaching the observed next states. If the starting state is itself probabilistic, its initial-state factor is included as well.

Prπ(trajectory) = Pr(S_t) × π(A_t | S_t) × Pr(S_{t+1} | S_t, A_t) × π(A_{t+1} | S_{t+1}) × Pr(S_{t+2} | S_{t+1}, A_{t+1}) × ...

thenenvironment transitionthenenvironment transitioncontinueInitial statePr(S_t)Observed actionπ(A_t | S_t)Next statePr(S_{t+1} | S_t, A_t)Next actionπ(A_{t+1} | S_{t+1})Following statetransition factorEndpoint
What factors make up the probability of a complete state-action trajectory, and how do the initial-state, transition, and policy probabilities combine over time?

Comparing Target and Behavior Likelihoods

The importance-sampling ratio is the probability of the observed trajectory under the target policy divided by the probability of that same trajectory under the behavior policy.

ρ = Prπ(trajectory) / Prb(trajectory)

A Two-Action Ratio

For two observed action selections, the target policy assigns probabilities 0.8 and 0.2. The behavior policy assigns probabilities 0.4 and 0.5 to those same observed actions. Calculate the policy-probability part of the importance-sampling ratio.

Write the target factors: The target-policy contribution is 0.8 × 0.2.

Write the behavior factors: The behavior-policy contribution is 0.4 × 0.5.

Form the ratio: Divide the target contribution by the behavior contribution: (0.8 × 0.2) / (0.4 × 0.5).

Evaluate: The numerator is 0.16 and the denominator is 0.20, so the ratio is 0.8.

The policy-probability ratio is 0.8. This observed action sequence is relatively less likely under the target policy than under the behavior policy, so its influence is reduced.

multiplyevaluatecontinuemultiplyRatio1First actiontarget ÷ behaviorPartial ratio0.8 ÷ 0.4 = 2Second actiontarget ÷ behaviorTrajectory ratio2 × (0.2 ÷ 0.5) = 0.8
How does the importance-sampling ratio accumulate as each action in a sampled trajectory contributes a target-policy probability divided by a behavior-policy probability?

Why Environment Factors Cancel

The target and behavior probabilities describe the same observed trajectory in the same MDP. Therefore, the initial-state and MDP transition factors appear in both complete trajectory probabilities. When the target probability is divided by the behavior probability, matching environment factors cancel. The policy-action factors differ because the policies differ, so those are the factors that remain in the final ratio.

ρ = [initial-state factor × target-action factors × transition factors] / [initial-state factor × behavior-action factors × transition factors] = target-action factors / behavior-action factors

divide by behavior trajectorymatching initial and transition factors cancelTarget trajectoryinitial × target actions ×transitionsPolicy-action ratiotarget actions ÷ behavioractionsBehavior trajectoryinitial × behavior actions× transitions
When the target-policy trajectory probability is divided by the behavior-policy trajectory probability, which factors cancel and why do only policy-action probabilities remain?

Using the Ratio for Prediction

In off-policy learning, returns from behavior-policy trajectories are weighted according to the relative probability of those trajectories under the target and behavior policies. A trajectory that is relatively more likely under the target policy receives greater influence. A trajectory that is relatively less likely under the target policy receives less influence. The importance-sampling ratio is therefore the bridge between the distribution that supplied the samples and the distribution whose expected value is being estimated.

Interpreting a Ratio

Interpret the ratio 0.8 from the two-action trajectory example.

Compare likelihoods: The target-policy probability for the observed action sequence is 0.8 times the behavior-policy probability for that sequence.

Adjust influence: Because the ratio is below one, the behavior-policy return from this trajectory receives less influence when estimating the target-policy quantity.

Connect to prediction: Applying this relative weighting allows behavior-policy samples to contribute to an estimate associated with the target policy.

The ratio is a reweighting factor: it indicates how strongly the observed behavior-policy trajectory should influence target-policy prediction.

observed probabilitycomparison probabilityreweights returnsBehavior policysupplies trajectoriesTarget policydefines quantityImportance-samplingratioTarget-policyestimate
Which policy supplies the experience, and which policy determines the quantity being estimated?

Practice Check

MEDIUM

A trajectory contains two observed actions. The target policy assigns probabilities 0.6 and 0.5 to them. The behavior policy assigns probabilities 0.3 and 0.5. Calculate the policy-action part of the importance-sampling ratio and state whether the trajectory receives greater or lesser influence for target-policy prediction.

Hints
  • Multiply the two target-policy probabilities.
  • Multiply the two behavior-policy probabilities.
  • Divide the target product by the behavior product.
  • Compare the result with 1.

What do you think happens?

Before calculating, predict whether the ratio will be greater than, equal to, or less than 1.

  • Greater than 1
  • Equal to 1
  • Less than 1
Reveal answer

Answer: Greater than 1

The target product is 0.6 × 0.5 = 0.30, while the behavior product is 0.3 × 0.5 = 0.15. Their ratio is 0.30 / 0.15 = 2, so the trajectory is relatively more likely under the target policy and receives greater influence.

Mistakes to Avoid

  • Using only the target-policy probability.

    The ratio compares target-policy likelihood with behavior-policy likelihood, so both probabilities are required.

    Fix: Divide the target-policy trajectory probability by the behavior-policy trajectory probability.

  • Keeping the MDP transition probabilities in the final ratio.

    The matching transition factors occur in both complete trajectory probabilities and cancel.

    Fix: Show the complete probabilities first, then cancel the common initial-state and transition factors.

  • Assuming the behavior policy must estimate its own value.

    Off-policy prediction specifically uses behavior-policy samples to estimate a quantity associated with a target policy.

    Fix: Keep the roles distinct: the behavior policy supplies experience, while the target policy defines the quantity being estimated.

  • Comparing different trajectories in the numerator and denominator.

    The importance-sampling ratio compares the likelihood of the same observed trajectory under two policies.

    Fix: Use identical observed states and actions in both trajectory probabilities.

When calculating a ratio, write the full trajectory structure before simplifying it. This makes it clear which factors come from action selection, which come from the MDP, and why the shared environment factors cancel.

Key Takeaways

  1. Importance sampling lets behavior-policy samples support estimates for a target policy.
  2. A trajectory probability is built from an initial-state factor, policy probabilities for observed actions, and MDP transition probabilities for observed state changes.
  3. The importance-sampling ratio is the target-policy trajectory probability divided by the behavior-policy trajectory probability.
  4. Because the same trajectory is evaluated in the same MDP, matching initial-state and transition factors cancel.
  5. The remaining policy-action ratio determines whether the trajectory receives greater or lesser influence in target-policy prediction.

Key Takeaways

  • Importance sampling addresses the mismatch between the policy that generates samples and the policy whose value is being estimated.
  • A state-action trajectory probability combines policy action-selection factors with MDP transition factors.
  • The importance-sampling ratio compares the same trajectory under the target and behavior policies.
  • Common initial-state and transition factors cancel, leaving a ratio of policy-action probabilities.
  • This ratio reweights behavior-policy returns so they can contribute to target-policy prediction.