Concepts / Truncated Importance Sampling Ratio

Truncated Importance Sampling Ratio

Discounting-aware importance sampling uses the structure of discounted returns instead of treating each return as an indivisible whole.

  • Programming

The Mismatch in Full-Episode Correction

A common first impression is that an entire return must always be multiplied by an importance sampling ratio covering the whole episode. That approach treats the return as one indivisible object. Discounting-aware importance sampling starts with a more precise question: which actions could have influenced the part of the return being estimated? If the relevant return becomes determined at an earlier horizon, policy-probability factors from later actions are not needed for that partial return. Including them can add variability without adding useful information.

The central alignment rule is simple: a partial return ending at horizon h should use an importance sampling ratio that also ends at horizon h.

multiplycontinue through hstopdo not includeρ at tincludedρ at t+1includedρ at t+hincludedhorizon hlater ratiosexcluded
Which action-probability factors belong to a return that ends at horizon h, and which later factors are excluded?

Tracing a Return by Horizon

Think of a discounted return as being assembled from several partial views of the future. One component may stop early, while another extends farther. Discounting can be viewed as partial termination at several possible horizons: later portions contribute less because the return structure gives them less weight. This interpretation makes it natural to ask what would happen if each component were estimated separately.

may stopmay continuemay continuestart statehorizon 1partial terminationhorizon 2partial terminationlater horizonsmaller contribution
How can discounting be understood as allowing a return to end at several possible horizons?

This does not mean that the physical episode literally ends at every possible horizon. It is an interpretation of the return's structure. Each horizon identifies a portion that can be considered on its own, and the discount rate determines how strongly different horizons matter. When the discount rate is smaller than 1, recognizing these different horizons becomes especially important.

Building a Flat Partial Return

A flat partial return is a nondiscounted sum that stops at a selected horizon. For a starting time t and horizon h, write its structure as G_h,t = reward at t + reward at t+1 + ... + reward at t+h. The important features are that the rewards are not discounted within this component and that the sum ends at the chosen cutoff. The full discounted return can then be understood through components with different horizons rather than as one indivisible object.

A Three-Step Flat Partial Return

Choose a horizon of three steps starting at time t.

Select the cutoff: Set h to 3. The partial return will stop after the reward associated with the third step in the selected range.

List the rewards: Collect the reward at t, the reward at t+1, and the reward at t+2, using the source's convention that the partial return contains rewards only through its chosen horizon.

Keep the sum flat: Add those rewards without discounting them inside this component.

Write the result: The resulting structure is G_3,t = reward at t + reward at t+1 + reward at t+2.

G_3,t is a nondiscounted partial sum containing only the selected three-step reward range.

addaddcontinue to hstopstarting point treward at tincludedreward at t+1includedreward at t+hincludedhorizon hstop sum
Which rewards belong to a flat partial return of horizon h, and where does the sum stop?

Matching the Ratio Cutoff

Once a return is decomposed into flat partial returns, each component should use only the policy-probability information relevant to its own horizon. For a flat partial return ending at h, the corresponding truncated ratio is formed from the target-policy-to-behavior-policy probability ratios beginning at time t and continuing through time t+h. In symbolic form, ρ_h,t is the product of π(A_k|S_k) divided by µ(A_k|S_k) for the relevant k values from t through t+h. The product stops where the partial return stops.

mismatched cutoffmatched cutoffpartial returnrewards through hpartial returnrewards through hfull ratiocontinues latertruncated ratiofactors through h
How do the reward terms and likelihood-ratio factors line up across the same sequence of time steps?

Pairing a Horizon-Two Return with Its Ratio

Construct the importance sampling correction for a flat partial return with horizon h equal to 2.

Identify the return range: The partial return contains only the rewards through the selected horizon 2.

Identify the ratio factors: Use the target-to-behavior probability ratios at the corresponding action and state pairs beginning at t and continuing through t+2.

Stop at the same point: Do not extend the product with later action-probability factors, because those factors belong beyond this partial return's horizon.

The horizon-two component is paired with a ratio whose product ends at t+2.

Why Later Factors Can Hurt

Suppose a particular discounted-return component is already determined by an early portion of a trajectory. A full-episode ratio still multiplies in action-probability factors from later steps. Those factors are unrelated to that component's horizon. They can increase variance without changing the expected update in the relevant setting. Truncating the ratio removes these unnecessary later factors while preserving the policy information needed to scale the partial return.

requiresfull episode continuesfull episode continuesreturn componentends at hratios through hneededlater rationot neededlater rationot needed
How can a full-episode product include action-ratio factors after the portion of the trajectory that contributes to a particular discounted return?

This is the practical reason to avoid treating every return as indivisible. A full-episode ratio may be mathematically associated with the entire trajectory, but that does not make every factor relevant to every return component. Discounting-aware estimation respects the different horizons created by the discounted-return structure.

Estimator viewReturn treatmentRatio treatment when the discount rate is less than 1
Conventional importance samplingTreats the return as one indivisible wholeUses a ratio covering the whole episode
Discounting-aware importance samplingUses the structure of discounted returns as partial returns with different horizonsTruncates each ratio to the horizon of its corresponding partial return

Applying the Design Rule

MEDIUM

A return component contains rewards only through horizon h. Decide whether its associated importance sampling ratio should include policy-probability factors after h. Then explain why.

Hints
  • Compare the endpoint of the reward sum with the endpoint of the probability-ratio product.
  • Ask whether a later action could influence the partial return being estimated.
  • Use the rule that the ratio stops where the partial return stops.

What do you think happens?

A flat partial return ends at horizon h. Should its associated ratio continue multiplying factors from times after h?

  • Yes, because every return must use the whole episode
  • No, because the ratio should match the partial return's horizon
  • Only if the partial return is nondiscounted
Reveal answer

Answer: No, because the ratio should match the partial return's horizon.

The partial return contains rewards only through its selected horizon. Continuing the ratio farther adds policy-probability factors that are not required for that component and can increase variance without changing the expected update in the relevant setting.

  • Using one full-episode ratio for every return component.

    The later factors are not required to scale that partial return and can add variability without useful information.

    Fix: Construct the ratio only through the same horizon as the partial return.

  • Treating a discounted return as one indivisible object.

    Discounting-aware importance sampling uses the structure of discounted returns rather than treating the whole return as a single unit.

    Fix: Decompose the return into flat partial returns and match each component with its own truncated ratio.

  • Discounting the terms inside a flat partial return.

    A flat partial return is defined as a nondiscounted sum that stops at a selected horizon.

    Fix: Keep the component flat, then use the broader discounted-return structure to determine how its horizons matter.

Operational Checklist

  1. Choose the return component's horizon h.
  2. Write the flat partial return as a nondiscounted sum that stops at h.
  3. Identify the target-to-behavior policy-probability ratios from the starting time through h.
  4. Multiply only those ratios into the component.
  5. Check that the reward endpoint and ratio endpoint are identical.
  6. When the discount rate is less than 1, consider the different horizons instead of automatically applying one full-episode correction.

Key Takeaways

  1. Discounting can be interpreted as partial termination at several possible horizons.
  2. A flat partial return is a nondiscounted reward sum that stops at a selected horizon.
  3. The importance sampling ratio for that component should stop at the same horizon.
  4. Later full-episode ratio factors may be unrelated to the component and can increase variance without adding useful information.
  5. Discounting-aware importance sampling differs from conventional importance sampling by respecting the structure and separate horizons of discounted returns.

Key Takeaways

  • Discounting creates return components with different effective horizons.
  • Flat partial returns are nondiscounted sums stopped at chosen horizons.
  • Each partial return should be paired with a truncated target-to-behavior probability ratio ending at the same horizon.
  • Full-episode factors after that horizon may add variance without contributing useful information to the component.
  • This horizon matching is the defining distinction between discounting-aware and conventional importance sampling in this setting.