Off-policy learning
Semi-gradient off-policy TD(0) extends the corresponding on-policy TD(0) algorithm by adding ρ_t.
Two policies, one learning problem
Off-policy prediction begins when the policy generating experience is different from the policy whose value function you want to learn. The data-generating policy is called the behavior policy. The policy being evaluated is called the target policy. Off-policy learning connects these two roles instead of requiring them to be the same.
Off-policy prediction is the task of learning the value function of a target policy from experience generated by a different behavior policy.
The extra correction in TD(0)
Semi-gradient off-policy TD(0) follows the corresponding on-policy TD(0) algorithm, but adds ρ_t. This additional term is the defining off-policy addition in the source description. The method remains a one-step, state-value algorithm; the key change is that the update includes this probability-based adjustment.
When comparing the on-policy and off-policy versions, look first for ρ_t. Its presence is the algorithmic addition that marks the off-policy version.
A traced decision
Following one off-policy experience
A behavior policy generates an experience by selecting an action. A different target policy is the policy whose value function is being learned. Identify the roles of the two policies and the purpose of ρ_t.
Identify the data source: The policy that selected the action and generated the experience is the behavior policy.
Identify the value being learned: The policy whose value function is being learned is the target policy.
Adjust the off-policy update: Because the policies are different, the semi-gradient off-policy TD(0) method includes ρ_t. This is the probability-based adjustment associated with the off-policy setting.
Keep the ideas separate: ρ_t identifies the off-policy extension. The definition of δ_t must still be selected according to whether the problem is episodic or continuing.
The experience comes from the behavior policy, the value function belongs to the target policy, and ρ_t is the defining additional term in the off-policy TD(0) update.
Episodic and continuing TD errors
The off-policy addition and the TD-error definition answer different questions. ρ_t answers how the method extends TD(0) to the off-policy setting. δ_t answers which TD-error definition fits the task. For an episodic problem, δ_t is defined for an episodic setting that may be discounted. For a continuing problem, δ_t is defined for an undiscounted setting using average reward.
| Question | Episodic setting | Continuing setting |
|---|---|---|
| What kind of task is described? | Episodic | Continuing |
| How is δ_t characterized? | Appropriate to a potentially discounted setting | Appropriate to an undiscounted average-reward setting |
| What determines the choice? | The task is episodic | The task is continuing |
The task setting determines the definition of δ_t; it does not replace the off-policy addition ρ_t.
Importance sampling the returns
When returns are observed under one policy but used to learn about another, importance sampling supplies a probability-based adjustment. It uses action-probability ratios to adjust observed returns so that off-policy experience can support prediction about the target policy.
| Estimator | How it averages | Statistical property | Practical implication |
|---|---|---|---|
| Ordinary importance sampling | Simple average of weighted returns | Unbiased, but can have large or infinite variance | May be less practical when variance is a concern |
| Weighted importance sampling | Weighted average of returns | Finite variance | Preferred in practice |
The trade-off is important. Ordinary importance sampling preserves unbiasedness, but its variance can be large or even infinite. Weighted importance sampling has finite variance and is preferred in practice, although the source describes the comparison as a trade-off between lower variance and bias.
Choosing the correct description
Treating off-policy learning as if the behavior and target policies must be identical.
Off-policy prediction is specifically about separating the policy generating experience from the policy whose value function is learned.
Fix:
Name both roles: behavior policy for the data and target policy for the value function.Thinking that ρ_t is the definition of the TD error.
ρ_t is the defining off-policy addition, while δ_t remains the TD-error term.
Fix:
Describe ρ_t as the off-policy adjustment and choose the definition of δ_t from the episodic or continuing task setting.Using the episodic description of δ_t for every problem.
Continuing problems use a definition appropriate to an undiscounted average-reward setting.
Fix:
First classify the task as episodic or continuing, then use the corresponding δ_t definition.Calling ordinary and weighted importance sampling interchangeable.
Ordinary sampling is unbiased but can have large or infinite variance, while weighted sampling has finite variance and is preferred in practice.
Fix:
Identify the averaging method and account for the variance and bias trade-off.
Use a three-question check: Are the behavior and target policies different? Is there a probability-based importance-sampling adjustment? Is the estimator ordinary or weighted? Then separately identify whether δ_t belongs to an episodic or continuing problem.
Practice the separation
A learning description says that one policy generates the experience, a different policy is being evaluated, and a weighted-average importance-sampling estimator is used. The task is continuing and undiscounted with average reward. Identify the behavior policy, the target policy, the off-policy addition, the appropriate δ_t setting, and the estimator choice.
Hints
- The policy generating experience is the behavior policy.
- The policy being evaluated is the target policy.
- The defining extra term in semi-gradient off-policy TD(0) is ρ_t.
- A continuing undiscounted average-reward task uses the continuing definition of δ_t.
- A weighted-average estimator corresponds to weighted importance sampling.
Practice answer
Classify each part of the description above.
Behavior policy: It is the policy that generates the experience.
Target policy: It is the different policy whose value function is being learned.
Off-policy addition: ρ_t is the additional term that extends the corresponding on-policy TD(0) algorithm.
TD-error setting: Because the task is continuing and undiscounted with average reward, δ_t uses the continuing-problem definition.
Estimator: A weighted-average procedure is weighted importance sampling, which has finite variance and is preferred in practice.
The description is an off-policy prediction problem using ρ_t, the continuing definition of δ_t, and weighted importance sampling.
Key takeaways
- Off-policy prediction learns the value function of a target policy from experience generated by a behavior policy.
- Semi-gradient off-policy TD(0) extends the corresponding on-policy method by adding ρ_t.
- The definition of δ_t is a separate, problem-dependent choice: episodic problems may be discounted, while continuing problems use an undiscounted average-reward setting.
- Importance sampling uses action-probability ratios to adjust off-policy returns.
- Ordinary importance sampling is unbiased but can have large or infinite variance; weighted importance sampling has finite variance and is preferred in practice.
Key Takeaways
- Off-policy learning separates the behavior policy that generates experience from the target policy being evaluated.
- The defining algorithmic addition in semi-gradient off-policy TD(0) is ρ_t.
- δ_t is still the TD-error term, but its definition depends on whether the problem is episodic or continuing.
- Importance sampling adjusts observed returns using action-probability ratios.
- Ordinary importance sampling is unbiased but may have high or infinite variance, whereas weighted importance sampling has finite variance and is preferred in practice.