Concepts / Off-policy learning

Off-policy learning

Semi-gradient off-policy TD(0) extends the corresponding on-policy TD(0) algorithm by adding ρ_t.

  • Programming

Two policies, one learning problem

Off-policy prediction begins when the policy generating experience is different from the policy whose value function you want to learn. The data-generating policy is called the behavior policy. The policy being evaluated is called the target policy. Off-policy learning connects these two roles instead of requiring them to be the same.

generatesused to learnBehavior policygenerates experienceExperienceobserved returnsTarget policyvalue function learned
How do the policy generating experience and the policy being evaluated differ, and how does experience move between them?

Off-policy prediction is the task of learning the value function of a target policy from experience generated by a different behavior policy.

The extra correction in TD(0)

Semi-gradient off-policy TD(0) follows the corresponding on-policy TD(0) algorithm, but adds ρ_t. This additional term is the defining off-policy addition in the source description. The method remains a one-step, state-value algorithm; the key change is that the update includes this probability-based adjustment.

identifiesdeterminesadded toused bySelected actionfrom behavior policyAction probabilitiesbehavior and targetρ_timportance adjustmentδ_tTD-error termTD(0) updateoff-policy version
How does the action selected by the behavior policy determine the off-policy adjustment, and where does ρ_t enter the TD(0) update?

When comparing the on-policy and off-policy versions, look first for ρ_t. Its presence is the algorithmic addition that marks the off-policy version.

A traced decision

Following one off-policy experience

A behavior policy generates an experience by selecting an action. A different target policy is the policy whose value function is being learned. Identify the roles of the two policies and the purpose of ρ_t.

Identify the data source: The policy that selected the action and generated the experience is the behavior policy.

Identify the value being learned: The policy whose value function is being learned is the target policy.

Adjust the off-policy update: Because the policies are different, the semi-gradient off-policy TD(0) method includes ρ_t. This is the probability-based adjustment associated with the off-policy setting.

Keep the ideas separate: ρ_t identifies the off-policy extension. The definition of δ_t must still be selected according to whether the problem is episodic or continuing.

The experience comes from the behavior policy, the value function belongs to the target policy, and ρ_t is the defining additional term in the off-policy TD(0) update.

Episodic and continuing TD errors

The off-policy addition and the TD-error definition answer different questions. ρ_t answers how the method extends TD(0) to the off-policy setting. δ_t answers which TD-error definition fits the task. For an episodic problem, δ_t is defined for an episodic setting that may be discounted. For a continuing problem, δ_t is defined for an undiscounted setting using average reward.

definesdefinesEpisodic problempotentially discountedδ_tepisodic definitionContinuing problemundiscounted average rewardδ_tcontinuing definition
Which terms and task assumptions belong to the TD error in an episodic problem versus a continuing problem?
QuestionEpisodic settingContinuing setting
What kind of task is described?EpisodicContinuing
How is δ_t characterized?Appropriate to a potentially discounted settingAppropriate to an undiscounted average-reward setting
What determines the choice?The task is episodicThe task is continuing

The task setting determines the definition of δ_t; it does not replace the off-policy addition ρ_t.

determinesextendssuppliesρ_toff-policy additionδ_ttask-dependent definitionTask settingepisodic or continuingOff-policy TD(0)one-step state-value method
Which part is the general off-policy addition, and which part changes because the task is episodic or continuing?

Importance sampling the returns

When returns are observed under one policy but used to learn about another, importance sampling supplies a probability-based adjustment. It uses action-probability ratios to adjust observed returns so that off-policy experience can support prediction about the target policy.

averaged simplyaveraged with weightsproducesproducesOff-policy returnsprobability-adjustedOrdinary samplingsimple averageOrdinary estimateunbiasedWeighted samplingweighted averageWeighted estimatefinite variance
How are ordinary and weighted estimates formed from the same off-policy returns, and where does normalization occur?
EstimatorHow it averagesStatistical propertyPractical implication
Ordinary importance samplingSimple average of weighted returnsUnbiased, but can have large or infinite varianceMay be less practical when variance is a concern
Weighted importance samplingWeighted average of returnsFinite variancePreferred in practice

The trade-off is important. Ordinary importance sampling preserves unbiasedness, but its variance can be large or even infinite. Weighted importance sampling has finite variance and is preferred in practice, although the source describes the comparison as a trade-off between lower variance and bias.

can havehasOrdinary samplingunbiasedLarge variancepossibly infiniteWeighted samplingbiased trade-offFinite variancepreferred in practice
Why can ordinary importance sampling be statistically attractive but practically difficult, while weighted importance sampling is preferred in practice?

Choosing the correct description

  • Treating off-policy learning as if the behavior and target policies must be identical.

    Off-policy prediction is specifically about separating the policy generating experience from the policy whose value function is learned.

    Fix: Name both roles: behavior policy for the data and target policy for the value function.

  • Thinking that ρ_t is the definition of the TD error.

    ρ_t is the defining off-policy addition, while δ_t remains the TD-error term.

    Fix: Describe ρ_t as the off-policy adjustment and choose the definition of δ_t from the episodic or continuing task setting.

  • Using the episodic description of δ_t for every problem.

    Continuing problems use a definition appropriate to an undiscounted average-reward setting.

    Fix: First classify the task as episodic or continuing, then use the corresponding δ_t definition.

  • Calling ordinary and weighted importance sampling interchangeable.

    Ordinary sampling is unbiased but can have large or infinite variance, while weighted sampling has finite variance and is preferred in practice.

    Fix: Identify the averaging method and account for the variance and bias trade-off.

Use a three-question check: Are the behavior and target policies different? Is there a probability-based importance-sampling adjustment? Is the estimator ordinary or weighted? Then separately identify whether δ_t belongs to an episodic or continuing problem.

Practice the separation

MEDIUM

A learning description says that one policy generates the experience, a different policy is being evaluated, and a weighted-average importance-sampling estimator is used. The task is continuing and undiscounted with average reward. Identify the behavior policy, the target policy, the off-policy addition, the appropriate δ_t setting, and the estimator choice.

Hints
  • The policy generating experience is the behavior policy.
  • The policy being evaluated is the target policy.
  • The defining extra term in semi-gradient off-policy TD(0) is ρ_t.
  • A continuing undiscounted average-reward task uses the continuing definition of δ_t.
  • A weighted-average estimator corresponds to weighted importance sampling.

Practice answer

Classify each part of the description above.

Behavior policy: It is the policy that generates the experience.

Target policy: It is the different policy whose value function is being learned.

Off-policy addition: ρ_t is the additional term that extends the corresponding on-policy TD(0) algorithm.

TD-error setting: Because the task is continuing and undiscounted with average reward, δ_t uses the continuing-problem definition.

Estimator: A weighted-average procedure is weighted importance sampling, which has finite variance and is preferred in practice.

The description is an off-policy prediction problem using ρ_t, the continuing definition of δ_t, and weighted importance sampling.

Key takeaways

  1. Off-policy prediction learns the value function of a target policy from experience generated by a behavior policy.
  2. Semi-gradient off-policy TD(0) extends the corresponding on-policy method by adding ρ_t.
  3. The definition of δ_t is a separate, problem-dependent choice: episodic problems may be discounted, while continuing problems use an undiscounted average-reward setting.
  4. Importance sampling uses action-probability ratios to adjust off-policy returns.
  5. Ordinary importance sampling is unbiased but can have large or infinite variance; weighted importance sampling has finite variance and is preferred in practice.

Key Takeaways

  • Off-policy learning separates the behavior policy that generates experience from the target policy being evaluated.
  • The defining algorithmic addition in semi-gradient off-policy TD(0) is ρ_t.
  • δ_t is still the TD-error term, but its definition depends on whether the problem is episodic or continuing.
  • Importance sampling adjusts observed returns using action-probability ratios.
  • Ordinary importance sampling is unbiased but may have high or infinite variance, whereas weighted importance sampling has finite variance and is preferred in practice.