Concepts / Exploration and Off-Policy Learning

Exploration and Off-Policy Learning

Off-policy prediction separates the policy generating experience from the policy whose value function is learned.

  • Programming

One Experience, Two Policies

Suppose the experience you have was produced by one policy, but the value function you want belongs to another. This is the central situation in off-policy prediction. The policy that generates experience is called the behavior policy. The policy whose value function is being learned is called the target policy.

off-policy prediction connects their rolesBehavior policygenerates experienceTarget policyvalue function learned
Which policy chooses each action, and which policy's value function is being learned?

The word off-policy refers to the separation between the policy generating the data and the policy being evaluated. If those policies are different, the setting is off-policy prediction.

Tracing the Roles Through Experience

Start with a trajectory, meaning an observed sequence of experience. That trajectory was produced according to the behavior policy. However, the learning goal concerns the target policy. The observed return therefore comes from behavior-policy experience, while the desired value concerns target-policy behavior.

inspect action probabilitiescompare policiesadjust observed returnsObserved trajectorygenerated by behaviorpolicyAction probabilitiesbehavior and targetProbability ratiosimportance-samplingadjustmentTarget value estimateestimated from adjustedreturns
How does a trajectory generated by the behavior policy get reweighted so it can estimate the target policy's value?

Importance sampling supplies the connection between the two policies. It uses action-probability ratios to adjust observed returns. The adjustment makes the return useful for estimating the value associated with the target policy rather than simply reporting the value of the behavior policy that generated the trajectory.

A Symbolic Reweighting Example

One Trajectory, Two Policy Assignments

A trajectory was generated by a behavior policy, but the value function of a different target policy is required. Explain the role of importance sampling without changing the observed trajectory.

Identify the source of the data: The behavior policy generated the trajectory and therefore determined which actions appeared in the observed experience.

Identify the evaluation goal: The target policy is the policy whose value function is being learned.

Compare action probabilities: Importance sampling uses ratios involving the target and behavior action probabilities for the observed actions.

Adjust the return: The observed return is reweighted using the probability adjustment so it can contribute to estimating the target policy's value.

The trajectory remains behavior-policy experience, but its contribution is adjusted for target-policy evaluation.

observed actionsnumerator informationdenominator informationmultiply return by adjustmentBehavior experienceobserved returnTarget probabilitiesfor observed actionsBehaviorprobabilitiesfor observed actionsProbability ratiosreweighting factorAdjusted returnsupports target evaluation
How does a trajectory generated by one policy become evidence about another policy's value?

Ordinary Importance Sampling

Ordinary importance sampling applies the trajectory's cumulative likelihood ratio to the return and then uses a simple average across the weighted returns. In other words, each observed return is adjusted according to the action-probability ratios along its trajectory, and the adjusted returns are averaged in the ordinary way.

multiplymultiplyaverage across trajectoriesObserved returnCumulativelikelihood ratioWeighted returnreturn adjusted for targetpolicySimple averageordinary estimator
How is a return multiplied by its trajectory's cumulative likelihood ratio to produce an ordinary importance-sampling estimate?

The important property of ordinary importance sampling is that it is unbiased. Its weakness is that it can have large or even infinite variance.

What do you think happens?

If the cumulative likelihood ratio is unusually large for one trajectory, what can happen to that trajectory's contribution in ordinary importance sampling?

  • It can be strongly amplified
  • It is automatically discarded
  • It has no effect on the estimate
Reveal answer

Answer: It can be strongly amplified

Ordinary importance sampling multiplies the return by the cumulative likelihood ratio. Large ratios can therefore produce large weighted returns, contributing to large or infinite variance.

Weighted Importance Sampling

Weighted importance sampling uses the same general idea of adjusting returns with importance weights, but it combines multiple returns with a weighted average. The weights are used in the averaging process rather than simply taking the ordinary average of the already weighted returns.

combine with weightsscale returnsadd weightsnormalizenormalizeOff-policy returnsmultiple trajectoriesImportance weightsone for each trajectoryWeighted return sumWeight sumWeighted averageweighted estimator
How are multiple off-policy returns normalized by the sum of their importance weights in weighted importance sampling?

The source distinguishes the two averaging methods directly: ordinary importance sampling uses a simple average of weighted returns, while weighted importance sampling uses a weighted average. Weighted importance sampling has finite variance and is preferred in practice.

Choosing Between the Estimators

variance and practical-use trade-offOrdinary importancesamplingunbiased; variance can belarge or infiniteWeighted importancesamplingfinite variance; preferredin practice
How do ordinary and weighted importance sampling differ in unbiasedness, variance, stability, and practical usefulness?
EstimatorHow returns are combinedVariance factPractical note
Ordinary importance samplingSimple average of weighted returnsUnbiased, but variance can be large or infiniteIts unbiasedness does not prevent unstable estimates
Weighted importance samplingWeighted average of returnsFinite variancePreferred in practice

The practical trade-off is therefore not simply about whether an estimator uses importance weights. Both do. The key choice is between ordinary averaging, which preserves unbiasedness but may have very large or infinite variance, and weighted averaging, which has finite variance and is preferred in practice.

Exploration as Data Generation

Exploration and off-policy learning fit together through the behavior policy. The behavior policy generates the experience. Off-policy prediction then uses that experience to learn about the target policy, while importance sampling adjusts the observed returns according to action-probability ratios.

generatesprovides returns and actionssupports predictionBehavior policygenerates experienceExperienceobserved trajectoriesImportance adjustmentaction-probability ratiosTarget value functionlearned from adjustedreturns
How can a behavior policy generate experience that supports learning about a different target policy?

Off-policy learning does not require the data-generating policy and the evaluated policy to be the same. That separation is the defining feature of off-policy prediction.

Mistakes in Policy Identification

  • Treating the behavior policy as the policy whose value is being learned

    The behavior policy generates experience, while the target policy is the one whose value function is learned.

    Fix: Label the two roles separately before applying any estimator.

  • Calling every return from behavior-policy experience a target-policy return without adjustment

    Off-policy prediction requires a probability-based adjustment to connect the observed experience to the target policy.

    Fix: Look for importance sampling and its action-probability ratios.

  • Assuming ordinary importance sampling is automatically stable because it is unbiased

    Ordinary importance sampling is unbiased but can have large or infinite variance.

    Fix: Consider weighted importance sampling when practical use and finite variance matter.

  • Confusing the names of the two estimators

    The distinction concerns the averaging method: ordinary importance sampling uses a simple average of weighted returns, while weighted importance sampling uses a weighted average.

    Fix: Ask how multiple adjusted returns are combined.

Check Your Understanding

MEDIUM

A data set was generated by policy A, but the value function of policy B is being learned. Identify the behavior policy and the target policy. Then state why an importance-sampling adjustment is needed and which estimator is preferred in practice when finite variance is important.

Hints
  • The behavior policy is defined by which policy generated the experience.
  • The target policy is defined by which policy's value function is being learned.
  • Importance sampling uses action-probability ratios to adjust observed returns.
  • Weighted importance sampling is preferred in practice because its variance is finite.

Answer Check

Resolve the policy roles and estimator choice in the practice scenario.

Behavior policy: Policy A is the behavior policy because it generated the data.

Target policy: Policy B is the target policy because its value function is being learned.

Importance sampling: The policies are different, so action-probability ratios are used to adjust the observed returns.

Estimator choice: Weighted importance sampling is preferred in practice because it has finite variance.

A is the behavior policy, B is the target policy, and weighted importance sampling is the practical choice when finite variance is important.

Summary

  1. Off-policy prediction separates the policy that generates experience from the policy whose value function is learned. Importance sampling uses action-probability ratios to adjust observed returns for this difference. Ordinary importance sampling uses a simple average of weighted returns and is unbiased, but its variance can be large or infinite. Weighted importance sampling uses a weighted average, has finite variance, and is preferred in practice.

Key Takeaways

  • The behavior policy generates experience; the target policy is the one being evaluated.
  • Off-policy prediction applies when these two policies are different.
  • Importance sampling uses action-probability ratios to adjust observed returns.
  • Ordinary importance sampling is unbiased but can have large or infinite variance.
  • Weighted importance sampling has finite variance and is preferred in practice.