Concepts / Interventions and Counterfactuals in Probabilistic Graphical Models

Interventions and Counterfactuals in Probabilistic Graphical Models

Off-policy learning separates the policy that gathers experiences from the policy being learned.

  • Programming

The Data-Policy Mismatch

Imagine that you want to learn about one policy, but the experiences available to you were collected while following another. This is the central situation in off-policy learning. The policy that gathers experiences is separate from the policy being learned, so the data source and the behavior or distribution of interest do not automatically match.

producesused to learn aboutBehavior policycollects experiencesExperiencesavailable dataTarget policybeing learned
How do experiences flow from the policy that collects data to a different policy being evaluated or learned?

The word off-policy describes how data collection and learning are separated: experiences come from one policy, while learning concerns another.

Tracing the Distribution Problem

Learning from a Different Experience Source

Suppose a learner wants to understand a target policy, but the available experiences were gathered under a behavior policy.

Collect: The behavior policy produces the experiences that are available for learning.

Compare: The learner notices that the source distribution of those experiences is not necessarily the target distribution associated with the policy being studied.

Correct: Importance sampling is relevant because it addresses the mismatch between samples from one distribution and the target distribution.

Learn: The corrected information is used to learn about the target policy rather than simply treating the source experiences as if they had been generated by it.

The key issue is not merely that data are available. The issue is whether the data-producing policy and the policy being learned are the same.

Importance sampling matters here because available samples may come from a different distribution than the distribution being studied. It provides the conceptual correction that lets samples from the source distribution contribute to reasoning about the target distribution.

generatesreceivesinformsSource policydata collectionSampled experiencesource distributionLikelihood ratioreweighting stepTarget estimatetarget distribution
How are samples generated by one policy reweighted so they can inform estimates under another policy?

Four Importance-Sampling Treatments

Importance sampling is not presented as one unchanging treatment in the source history. Ordinary importance sampling, weighted or normalized importance sampling, per-reward treatments, and discounting-aware treatments represent distinct points in the development of the method. The distinctions concern how the correction is applied across a trajectory and across the rewards or outcomes being considered.

Ordinarytrajectory treatmentWeighted ornormalizednormalized treatmentPer-rewardreward-level treatmentDiscounting-awarediscount-sensitivetreatment
How do the main treatments differ in where and how they apply correction weights?
TreatmentWhat the name highlightsWhat it helps the learner distinguish
Ordinary importance samplingThe basic correction for a source-target distribution mismatchThe general idea of reusing samples from a different distribution
Weighted or normalized importance samplingA weighted or normalized form of the correctionA distinct treatment from the ordinary form
Per-reward importance samplingCorrection considered at the level of individual rewardsHow weighting can be associated with separate rewards rather than only an entire trajectory
Discounting-aware importance samplingCorrection considered together with discountingHow the treatment changes when discounting is part of the learning setup

When reading or implementing an off-policy method, first ask which treatment is being used. The phrase importance sampling alone does not tell you whether the method is ordinary, weighted or normalized, per-reward, or discounting-aware.

Observation and Intervention

The topic connects off-policy learning with interventions and counterfactuals in probabilistic graphical models. A useful first distinction is between observing what occurred and considering a situation in which a variable or choice is externally set. Observation describes available evidence. Intervention represents a deliberate change to the situation being analyzed.

Observed variablerecords what occurredIntervened variableexternally set
What changes when a variable is externally set by an intervention instead of merely observed?

The connection to off-policy learning is conceptual: both settings separate what was observed or collected from the alternative condition or policy that we want to understand.

Counterfactual Policy Questions

What do you think happens?

A behavior policy produced an experience, but the learner is studying a different target policy. Can that experience still be relevant to the target policy?

  • No, because only experiences generated by the target policy can ever be used.
  • Yes, because off-policy learning is designed to learn from experiences collected under a different policy.
  • Only if the behavior policy and target policy are renamed to be the same.
Reveal answer

Answer: Yes, because off-policy learning separates the policy that gathers experiences from the policy being learned.

The distribution mismatch remains important, which is why importance sampling is relevant. The experience is not treated as if it came from the target policy without correction.

A counterfactual question asks about an alternative condition: given what was collected, what would be learned or expected under a different policy or action? In this article's setting, the alternative is represented by the target policy rather than the policy that generated the available experience.

reused to studysupports reasoning aboutCollectedexperiencebehavior policyAlternative policytarget policyCounterfactualoutcomewhat would be learned
How can the same collected experience support reasoning about what would have happened under a different policy?

From Causal Ideas to Later Methods

The source presents importance sampling in off-policy learning as part of a longer research history rather than as an isolated technique. That history connects off-policy learning with interventions and counterfactuals in probabilistic graphical models. It also leads toward combinations with temporal-difference learning, eligibility traces, and approximation methods.

connected ideashistorical connectioncombined withcombined withcombined withInterventionsprobabilistic graphicalmodelsCounterfactualsalternative conditionsOff-policy learningdifferent data and targetpoliciesTemporal-differencelearninglater combinationEligibility traceslater combinationApproximation methodslater combination
How are interventions and counterfactuals connected to off-policy learning and later reinforcement-learning methods?

The historical lesson is also a practical warning. Once off-policy correction is combined with temporal-difference learning, eligibility traces, or approximation methods, additional subtle issues arise. The basic source-target mismatch remains the starting point, but later methods add further structure that must be analyzed separately.

Mistakes in Reading Off-Policy Methods

  • Treating the behavior policy and target policy as the same thing.

    Off-policy learning is defined by separating the policy that gathers experiences from the policy being learned.

    Fix: Name both policies explicitly before analyzing the method.

  • Ignoring the distribution mismatch.

    Importance sampling is relevant precisely because available samples may come from a different distribution than the target distribution.

    Fix: Identify the source distribution, the target distribution, and the correction treatment.

  • Assuming every importance-sampling method has the same treatment.

    The source identifies these as distinct points in the development of the method.

    Fix: Check which treatment the method specifies and what level of the trajectory or reward it addresses.

  • Confusing an observed experience with a counterfactual experience.

    The alternative policy is the condition being studied, not necessarily the source of the collected data.

    Fix: Describe the collected experience and the alternative policy separately.

  • Assuming the historical connection ends with importance sampling.

    The source explicitly connects off-policy learning to combinations with these later methods.

    Fix: Treat basic off-policy correction as a foundation for later combinations, not as the complete history.

Check Your Understanding

MEDIUM

A dataset was collected under a behavior policy, but a researcher wants to learn about a target policy. Explain why this is an off-policy setting, why importance sampling may be relevant, and which details you would check before deciding whether the method uses ordinary, weighted or normalized, per-reward, or discounting-aware treatment.

Hints
  • Start by naming the policy that produced the experiences and the policy being learned.
  • State the distribution mismatch explicitly.
  • List the four importance-sampling treatments as distinct possibilities rather than assuming they are interchangeable.
MEDIUM

Describe the difference between an observed experience and a counterfactual question in this setting. Then explain how the same collected experience can be relevant to reasoning about a different policy without claiming that the different policy actually generated the experience.

Hints
  • Observation concerns what was collected.
  • A counterfactual concerns an alternative policy or condition.
  • Mention the need to account for the source-target mismatch.

The Essential Distinctions

  1. Off-policy learning uses experiences gathered by one policy to learn about another policy.
  2. The central technical problem is a mismatch between the distribution that produced the samples and the target distribution being studied.
  3. Importance sampling is relevant because it corrects for that source-target mismatch.
  4. Ordinary, weighted or normalized, per-reward, and discounting-aware importance sampling are distinct treatments in the method's development.
  5. The topic connects interventions and counterfactuals with later combinations involving temporal-difference learning, eligibility traces, and approximation methods.

Key Takeaways

  • Off-policy learning separates data collection from the policy being learned.
  • Importance sampling addresses the mismatch between the source distribution and the target distribution.
  • Ordinary, weighted or normalized, per-reward, and discounting-aware treatments should be distinguished.
  • Interventions and counterfactuals provide related ways to think about externally changed or alternative conditions.
  • Off-policy learning later connects with temporal-difference learning, eligibility traces, and approximation methods.