Interventions and Counterfactuals in Probabilistic Graphical Models
Off-policy learning separates the policy that gathers experiences from the policy being learned.
The Data-Policy Mismatch
Imagine that you want to learn about one policy, but the experiences available to you were collected while following another. This is the central situation in off-policy learning. The policy that gathers experiences is separate from the policy being learned, so the data source and the behavior or distribution of interest do not automatically match.
The word off-policy describes how data collection and learning are separated: experiences come from one policy, while learning concerns another.
Tracing the Distribution Problem
Learning from a Different Experience Source
Suppose a learner wants to understand a target policy, but the available experiences were gathered under a behavior policy.
Collect: The behavior policy produces the experiences that are available for learning.
Compare: The learner notices that the source distribution of those experiences is not necessarily the target distribution associated with the policy being studied.
Correct: Importance sampling is relevant because it addresses the mismatch between samples from one distribution and the target distribution.
Learn: The corrected information is used to learn about the target policy rather than simply treating the source experiences as if they had been generated by it.
The key issue is not merely that data are available. The issue is whether the data-producing policy and the policy being learned are the same.
Importance sampling matters here because available samples may come from a different distribution than the distribution being studied. It provides the conceptual correction that lets samples from the source distribution contribute to reasoning about the target distribution.
Four Importance-Sampling Treatments
Importance sampling is not presented as one unchanging treatment in the source history. Ordinary importance sampling, weighted or normalized importance sampling, per-reward treatments, and discounting-aware treatments represent distinct points in the development of the method. The distinctions concern how the correction is applied across a trajectory and across the rewards or outcomes being considered.
| Treatment | What the name highlights | What it helps the learner distinguish |
|---|---|---|
| Ordinary importance sampling | The basic correction for a source-target distribution mismatch | The general idea of reusing samples from a different distribution |
| Weighted or normalized importance sampling | A weighted or normalized form of the correction | A distinct treatment from the ordinary form |
| Per-reward importance sampling | Correction considered at the level of individual rewards | How weighting can be associated with separate rewards rather than only an entire trajectory |
| Discounting-aware importance sampling | Correction considered together with discounting | How the treatment changes when discounting is part of the learning setup |
When reading or implementing an off-policy method, first ask which treatment is being used. The phrase importance sampling alone does not tell you whether the method is ordinary, weighted or normalized, per-reward, or discounting-aware.
Observation and Intervention
The topic connects off-policy learning with interventions and counterfactuals in probabilistic graphical models. A useful first distinction is between observing what occurred and considering a situation in which a variable or choice is externally set. Observation describes available evidence. Intervention represents a deliberate change to the situation being analyzed.
The connection to off-policy learning is conceptual: both settings separate what was observed or collected from the alternative condition or policy that we want to understand.
Counterfactual Policy Questions
What do you think happens?
A behavior policy produced an experience, but the learner is studying a different target policy. Can that experience still be relevant to the target policy?
Reveal answer
Answer: Yes, because off-policy learning separates the policy that gathers experiences from the policy being learned.
The distribution mismatch remains important, which is why importance sampling is relevant. The experience is not treated as if it came from the target policy without correction.
A counterfactual question asks about an alternative condition: given what was collected, what would be learned or expected under a different policy or action? In this article's setting, the alternative is represented by the target policy rather than the policy that generated the available experience.
From Causal Ideas to Later Methods
The source presents importance sampling in off-policy learning as part of a longer research history rather than as an isolated technique. That history connects off-policy learning with interventions and counterfactuals in probabilistic graphical models. It also leads toward combinations with temporal-difference learning, eligibility traces, and approximation methods.
The historical lesson is also a practical warning. Once off-policy correction is combined with temporal-difference learning, eligibility traces, or approximation methods, additional subtle issues arise. The basic source-target mismatch remains the starting point, but later methods add further structure that must be analyzed separately.
Mistakes in Reading Off-Policy Methods
Treating the behavior policy and target policy as the same thing.
Off-policy learning is defined by separating the policy that gathers experiences from the policy being learned.
Fix:
Name both policies explicitly before analyzing the method.Ignoring the distribution mismatch.
Importance sampling is relevant precisely because available samples may come from a different distribution than the target distribution.
Fix:
Identify the source distribution, the target distribution, and the correction treatment.Assuming every importance-sampling method has the same treatment.
The source identifies these as distinct points in the development of the method.
Fix:
Check which treatment the method specifies and what level of the trajectory or reward it addresses.Confusing an observed experience with a counterfactual experience.
The alternative policy is the condition being studied, not necessarily the source of the collected data.
Fix:
Describe the collected experience and the alternative policy separately.Assuming the historical connection ends with importance sampling.
The source explicitly connects off-policy learning to combinations with these later methods.
Fix:
Treat basic off-policy correction as a foundation for later combinations, not as the complete history.
Check Your Understanding
A dataset was collected under a behavior policy, but a researcher wants to learn about a target policy. Explain why this is an off-policy setting, why importance sampling may be relevant, and which details you would check before deciding whether the method uses ordinary, weighted or normalized, per-reward, or discounting-aware treatment.
Hints
- Start by naming the policy that produced the experiences and the policy being learned.
- State the distribution mismatch explicitly.
- List the four importance-sampling treatments as distinct possibilities rather than assuming they are interchangeable.
Describe the difference between an observed experience and a counterfactual question in this setting. Then explain how the same collected experience can be relevant to reasoning about a different policy without claiming that the different policy actually generated the experience.
Hints
- Observation concerns what was collected.
- A counterfactual concerns an alternative policy or condition.
- Mention the need to account for the source-target mismatch.
The Essential Distinctions
- Off-policy learning uses experiences gathered by one policy to learn about another policy.
- The central technical problem is a mismatch between the distribution that produced the samples and the target distribution being studied.
- Importance sampling is relevant because it corrects for that source-target mismatch.
- Ordinary, weighted or normalized, per-reward, and discounting-aware importance sampling are distinct treatments in the method's development.
- The topic connects interventions and counterfactuals with later combinations involving temporal-difference learning, eligibility traces, and approximation methods.
Key Takeaways
- Off-policy learning separates data collection from the policy being learned.
- Importance sampling addresses the mismatch between the source distribution and the target distribution.
- Ordinary, weighted or normalized, per-reward, and discounting-aware treatments should be distinguished.
- Interventions and counterfactuals provide related ways to think about externally changed or alternative conditions.
- Off-policy learning later connects with temporal-difference learning, eligibility traces, and approximation methods.