On-policy Prediction
Off-policy learning separates the policy that generates data from the policy being evaluated.
One State, Two Policies
Suppose you want to estimate the value of one Blackjack state, but the episodes available to you were generated by a policy different from the policy you want to evaluate. This is the central situation addressed by off-policy estimation. One policy supplies the experience, while another policy is the subject of the value estimate.
Off-policy learning separates the policy that generates data from the policy being evaluated. The data-generating policy is called the behavior policy, and the policy whose value is estimated is called the target policy.
The Blackjack Evaluation
The example evaluates one particular Blackjack state: the dealer is showing a deuce, the player has a sum of 13, and the player has a usable ace. Every sampled episode begins from this state, and the learning process uses complete episodes to estimate the value of that state.
| Role | Policy in the example | Purpose |
|---|---|---|
| Behavior policy | Randomly chooses hit or stick | Generates the episodes |
| Target policy | Sticks only when the player's sum is 20 or 21 | Is evaluated |
The two policies have different roles, even though the same sampled episodes provide the data.
The important separation is therefore not between two Blackjack states. It is between two policies: the random policy supplies the experience, while the policy that sticks only on 20 or 21 is the policy whose value is estimated.
Reweighting Sampled Returns
An episode generated by the behavior policy is not automatically an unbiased description of what the target policy would produce, because the two policies make choices differently. Ordinary importance sampling addresses this difference by attaching an importance-sampling weight to the episode's return. The weight adjusts the contribution of that return so that the collection of behavior-policy episodes can estimate the target policy's value.
One behavior-policy episode
An episode begins at the specified Blackjack state, is generated by the random hit-or-stick behavior policy, and produces a return. How does ordinary importance sampling use it?
1. Generate: The behavior policy generates a complete episode from the state with a dealer showing a deuce, player sum 13, and a usable ace.
2. Observe: The episode supplies a return, meaning the outcome used for the Monte Carlo estimate.
3. Compare: The actions selected while generating the episode are compared with the action choices specified by the target policy, which sticks only on 20 or 21.
4. Reweight: Ordinary importance sampling uses the resulting importance-sampling weight to adjust the episode's return before combining it with other weighted returns.
The episode remains behavior-policy data, but its weighted contribution is used to estimate the target policy's value.
Normalizing the Evidence
Weighted importance sampling uses the same idea of weighting returns, but it normalizes the weighted returns by the sum of their importance weights. In other words, it does not simply combine weighted returns directly; it also accounts for the total weight attached to the available evidence.
| Method | How returns are used | Reported behavior |
|---|---|---|
| Ordinary importance sampling | Uses importance weights to combine returns | Can have high variance |
| Weighted importance sampling | Normalizes weighted returns by the sum of the importance weights | Typically has lower initial error |
Learning Across Episodes
The experiment began every run with estimates of zero and then learned from 10,000 episodes. The process was repeated in 100 independent runs. At each episode count, the researchers averaged the squared error across those runs. This measures how far the learned estimate was from the value being estimated as learning progressed.
Both algorithms' error approached zero as more episodes were used. Weighted importance sampling had much lower error at the beginning. The source identifies this lower initial error as typical in practice, while ordinary importance sampling can have high variance. The result is not that ordinary importance sampling fails; rather, the two methods can behave differently early in learning even though both improve toward the desired value.
Mistakes in Policy Roles
Treating the random hit-or-stick policy as the policy being evaluated.
In the example, the random policy is the behavior policy. The target policy is the policy that sticks only on 20 or 21.
Fix:
Ask two separate questions: Which policy generated the data? Which policy's value is being estimated?Assuming that importance sampling changes the sampled episodes into target-policy episodes.
The episode still came from the behavior policy. Importance sampling changes the episode's contribution to the estimate through a weight.
Fix:
Describe the data as behavior-policy data and the adjusted estimate as an estimate of the target policy's value.Saying that weighted importance sampling ignores importance weights.
Both methods use weighted returns. Weighted importance sampling additionally normalizes by the sum of the importance weights.
Fix:
Remember: ordinary importance sampling weights returns; weighted importance sampling weights and normalizes them.Interpreting lower initial error as permanent superiority.
The reported results say that both methods' error approached zero, while weighted importance sampling had much lower error at the beginning.
Fix:
Separate early learning behavior from the longer-run result reported for the experiment.
Check Your Understanding
For the Blackjack experiment, explain in your own words why an episode generated by the random hit-or-stick policy can still contribute to an estimate of the policy that sticks only on 20 or 21. Then state the difference between ordinary and weighted importance sampling.
Hints
- Start by naming the behavior policy and the target policy.
- Explain what the importance-sampling weight does to an episode return.
- Mention the sum of the importance weights when describing weighted importance sampling.
What do you think happens?
After many episodes in the reported experiment, what happened to the error of both ordinary and weighted importance sampling?
Reveal answer
Answer: Both errors approached zero
The reported experiment found that both algorithms' error approached zero. Weighted importance sampling had much lower error at the beginning.
Key Takeaways
- Off-policy estimation uses data from a behavior policy to estimate the value of a different target policy.
- The Blackjack state in the example has a dealer showing a deuce, a player sum of 13, and a usable ace.
- The random hit-or-stick policy is the behavior policy, while sticking only on 20 or 21 is the target policy.
- Ordinary importance sampling weights episode returns; weighted importance sampling also normalizes those weighted returns by the sum of the importance weights.
- In the reported experiment, both methods' error approached zero, while weighted importance sampling typically had lower initial error.
Key Takeaways
- Off-policy learning separates the policy that generates experience from the policy being evaluated.
- In the Blackjack example, random hit-or-stick generates the data, and sticking only on 20 or 21 is evaluated.
- Ordinary importance sampling reweights returns, whereas weighted importance sampling normalizes the weighted returns.
- Both methods' error approached zero in the reported experiment, but weighted importance sampling had much lower initial error.