Truncated Importance Sampling Ratio
Discounting-aware importance sampling uses the structure of discounted returns instead of treating each return as an indivisible whole.
The Mismatch in Full-Episode Correction
A common first impression is that an entire return must always be multiplied by an importance sampling ratio covering the whole episode. That approach treats the return as one indivisible object. Discounting-aware importance sampling starts with a more precise question: which actions could have influenced the part of the return being estimated? If the relevant return becomes determined at an earlier horizon, policy-probability factors from later actions are not needed for that partial return. Including them can add variability without adding useful information.
The central alignment rule is simple: a partial return ending at horizon h should use an importance sampling ratio that also ends at horizon h.
Tracing a Return by Horizon
Think of a discounted return as being assembled from several partial views of the future. One component may stop early, while another extends farther. Discounting can be viewed as partial termination at several possible horizons: later portions contribute less because the return structure gives them less weight. This interpretation makes it natural to ask what would happen if each component were estimated separately.
This does not mean that the physical episode literally ends at every possible horizon. It is an interpretation of the return's structure. Each horizon identifies a portion that can be considered on its own, and the discount rate determines how strongly different horizons matter. When the discount rate is smaller than 1, recognizing these different horizons becomes especially important.
Building a Flat Partial Return
A flat partial return is a nondiscounted sum that stops at a selected horizon. For a starting time t and horizon h, write its structure as G_h,t = reward at t + reward at t+1 + ... + reward at t+h. The important features are that the rewards are not discounted within this component and that the sum ends at the chosen cutoff. The full discounted return can then be understood through components with different horizons rather than as one indivisible object.
A Three-Step Flat Partial Return
Choose a horizon of three steps starting at time t.
Select the cutoff: Set h to 3. The partial return will stop after the reward associated with the third step in the selected range.
List the rewards: Collect the reward at t, the reward at t+1, and the reward at t+2, using the source's convention that the partial return contains rewards only through its chosen horizon.
Keep the sum flat: Add those rewards without discounting them inside this component.
Write the result: The resulting structure is G_3,t = reward at t + reward at t+1 + reward at t+2.
G_3,t is a nondiscounted partial sum containing only the selected three-step reward range.
Matching the Ratio Cutoff
Once a return is decomposed into flat partial returns, each component should use only the policy-probability information relevant to its own horizon. For a flat partial return ending at h, the corresponding truncated ratio is formed from the target-policy-to-behavior-policy probability ratios beginning at time t and continuing through time t+h. In symbolic form, ρ_h,t is the product of π(A_k|S_k) divided by µ(A_k|S_k) for the relevant k values from t through t+h. The product stops where the partial return stops.
Pairing a Horizon-Two Return with Its Ratio
Construct the importance sampling correction for a flat partial return with horizon h equal to 2.
Identify the return range: The partial return contains only the rewards through the selected horizon 2.
Identify the ratio factors: Use the target-to-behavior probability ratios at the corresponding action and state pairs beginning at t and continuing through t+2.
Stop at the same point: Do not extend the product with later action-probability factors, because those factors belong beyond this partial return's horizon.
The horizon-two component is paired with a ratio whose product ends at t+2.
Why Later Factors Can Hurt
Suppose a particular discounted-return component is already determined by an early portion of a trajectory. A full-episode ratio still multiplies in action-probability factors from later steps. Those factors are unrelated to that component's horizon. They can increase variance without changing the expected update in the relevant setting. Truncating the ratio removes these unnecessary later factors while preserving the policy information needed to scale the partial return.
This is the practical reason to avoid treating every return as indivisible. A full-episode ratio may be mathematically associated with the entire trajectory, but that does not make every factor relevant to every return component. Discounting-aware estimation respects the different horizons created by the discounted-return structure.
| Estimator view | Return treatment | Ratio treatment when the discount rate is less than 1 |
|---|---|---|
| Conventional importance sampling | Treats the return as one indivisible whole | Uses a ratio covering the whole episode |
| Discounting-aware importance sampling | Uses the structure of discounted returns as partial returns with different horizons | Truncates each ratio to the horizon of its corresponding partial return |
Applying the Design Rule
A return component contains rewards only through horizon h. Decide whether its associated importance sampling ratio should include policy-probability factors after h. Then explain why.
Hints
- Compare the endpoint of the reward sum with the endpoint of the probability-ratio product.
- Ask whether a later action could influence the partial return being estimated.
- Use the rule that the ratio stops where the partial return stops.
What do you think happens?
A flat partial return ends at horizon h. Should its associated ratio continue multiplying factors from times after h?
Reveal answer
Answer: No, because the ratio should match the partial return's horizon.
The partial return contains rewards only through its selected horizon. Continuing the ratio farther adds policy-probability factors that are not required for that component and can increase variance without changing the expected update in the relevant setting.
Using one full-episode ratio for every return component.
The later factors are not required to scale that partial return and can add variability without useful information.
Fix:
Construct the ratio only through the same horizon as the partial return.Treating a discounted return as one indivisible object.
Discounting-aware importance sampling uses the structure of discounted returns rather than treating the whole return as a single unit.
Fix:
Decompose the return into flat partial returns and match each component with its own truncated ratio.Discounting the terms inside a flat partial return.
A flat partial return is defined as a nondiscounted sum that stops at a selected horizon.
Fix:
Keep the component flat, then use the broader discounted-return structure to determine how its horizons matter.
Operational Checklist
- Choose the return component's horizon h.
- Write the flat partial return as a nondiscounted sum that stops at h.
- Identify the target-to-behavior policy-probability ratios from the starting time through h.
- Multiply only those ratios into the component.
- Check that the reward endpoint and ratio endpoint are identical.
- When the discount rate is less than 1, consider the different horizons instead of automatically applying one full-episode correction.
Key Takeaways
- Discounting can be interpreted as partial termination at several possible horizons.
- A flat partial return is a nondiscounted reward sum that stops at a selected horizon.
- The importance sampling ratio for that component should stop at the same horizon.
- Later full-episode ratio factors may be unrelated to the component and can increase variance without adding useful information.
- Discounting-aware importance sampling differs from conventional importance sampling by respecting the structure and separate horizons of discounted returns.
Key Takeaways
- Discounting creates return components with different effective horizons.
- Flat partial returns are nondiscounted sums stopped at chosen horizons.
- Each partial return should be paired with a truncated target-to-behavior probability ratio ending at the same horizon.
- Full-episode factors after that horizon may add variance without contributing useful information to the component.
- This horizon matching is the defining distinction between discounting-aware and conventional importance sampling in this setting.