Flat Partial Returns
Discounting-aware importance sampling uses the structure of discounted returns instead of treating each return as an indivisible whole.
The Whole-Return Habit
A common first impression is that an entire return must always be multiplied by an importance sampling ratio covering the whole episode. That approach treats the return as one indivisible object. Flat partial returns begin with a more precise question: which actions could actually have influenced the part of the return being estimated?
When discounting makes a return effectively determined at an earlier horizon, policy-probability factors from later actions are not needed for that partial return. Including those later factors can add variability without adding useful information. Discounting-aware importance sampling therefore uses the structure of the discounted return rather than treating the return as a single indivisible whole.
A Return Ends at Different Horizons
Discounting can be viewed as partial termination at several possible horizons. Instead of thinking only about one complete episode return, think of the discounted return as being built from components that stop at different points. A component that stops at horizon h contains rewards only through that horizon.
The diagram represents the construction rule: choose a horizon, include the rewards through that horizon, and stop there. The corresponding policy information should stop at the same point. The horizon is not an extra reward; it identifies where this partial component ends.
Building One Partial Component
A horizon-three flat partial return
Suppose a return begins at time t and the chosen horizon is h = 3. Construct the flat partial return and identify the policy-ratio information that belongs with it.
Choose the stopping point: The selected horizon is h = 3, so this component ends after the reward at t+3.
Collect the rewards: The flat partial return contains the nondiscounted reward sum from t through t+3. It does not include rewards after t+3.
Match the policy information: The associated importance sampling ratio contains the target-to-behavior policy-probability factors from time t through the same endpoint, t+3.
Reject later factors: Factors belonging to actions after t+3 are outside this partial return's horizon. They are not required to scale this component.
The horizon-three component is a flat reward sum through t+3 paired with a truncated ratio through t+3.
Using the source notation, the partial return is represented as G_h_t, and its matching truncated ratio is represented as rho_h_t. The ratio begins with the target-policy probability divided by the behavior-policy probability at time t, then continues with the same kind of target-to-behavior factors through time t+h. It stops where the partial return stops.
The index map shows that rho_h_t is not a full-episode product by default. It is the product of the relevant target-to-behavior probability ratios through t+h. The later factors remain part of the episode's later behavior, but they are not part of the ratio assigned to this partial component.
Discounting as Early Stopping
The partial-termination interpretation gives an intuitive way to understand discounting. At each possible point, the return may be viewed as continuing to a later horizon or terminating for the component being considered. The discount rate determines how strongly later horizons contribute, while each flat partial return remains an undiscounted sum up to its own stopping point.
This interpretation explains why a single full-episode ratio is poorly matched to every part of a discounted return. Different components have different stopping horizons. Each component should use only the action information that could have influenced it before it stopped.
Why Later Factors Are Irrelevant
| Estimator view | Return portion | Policy-ratio horizon | Main concern |
|---|---|---|---|
| Conventional full-return importance sampling | Treats the return as one indivisible whole | Whole episode | Later factors may be unrelated to an earlier return component |
| Discounting-aware flat partial returns | Separates the return into flat sums with selected horizons | Same horizon as the partial return | Avoids unnecessary later factors for that component |
A full-episode ratio can contain factors for actions taken after the partial return has already ended. Those factors are not required to scale the selected partial return. Continuing the product farther can increase variance without changing the expected update in the relevant setting.
Mistakes to Avoid
Multiplying every partial return by a ratio covering the whole episode.
The later factors are not required to scale that component and can add variability without useful information.
Fix:
Stop the ratio at the same horizon as the flat partial return.Treating a discounted return as one indivisible object.
Discounting can be viewed as partial termination at several possible horizons.
Fix:
Decompose the return conceptually into flat partial returns with their own stopping points.Discounting the flat partial return internally.
A flat partial return is a nondiscounted sum that stops at a selected horizon.
Fix:
Use the horizon to define the component; use the collection of horizons to represent the discounting structure.Assuming one ratio horizon is automatically appropriate for every component.
A ratio appropriate for one component need not extend through the entire episode or match another component's horizon.
Fix:
Match each component with policy-probability information relevant to its own horizon.
Check Your Construction
Choose a horizon h for a return beginning at time t. Describe which rewards belong in G_h_t, then describe where rho_h_t must stop. Finally, explain why a policy-ratio factor from an action after t+h is unnecessary for this component.
Hints
- Start by naming the final included reward: it is the reward at the selected horizon.
- The flat partial return includes rewards through that point and no later rewards.
- The matching ratio includes target-to-behavior policy-probability factors through the same point.
What do you think happens?
A partial return stops at t+h. Should its associated importance sampling ratio include a factor for an action after t+h?
Reveal answer
Answer: No, because the ratio should stop where the partial return stops.
The partial return contains rewards only through its chosen horizon, so its associated ratio needs policy-probability factors only through that same point. Later factors are not required for this component and can increase variance without changing the expected update in the relevant setting.
The Matching Principle
- A flat partial return is a nondiscounted reward sum that stops at a selected horizon.
- Discounting can be interpreted as partial termination across several possible horizons.
- The ratio paired with a partial return should contain target-to-behavior policy-probability factors only through that return's horizon.
- A full-episode ratio may include later factors unrelated to an earlier discounted-return component.
- When the discount rate is less than 1, discounting-aware estimation uses the different horizons of return components instead of applying one whole-episode ratio indiscriminately.
Key Takeaways
- Flat partial returns separate a discounted return into nondiscounted sums that stop at selected horizons.
- Discounting can be understood as partial termination at different possible horizons.
- The importance sampling ratio should be truncated to the same horizon as the partial return.
- Later policy-ratio factors can add variance when they correspond to actions outside the return component's effective horizon.
- Discounting-aware importance sampling therefore differs from conventional full-episode importance sampling when the discount rate is less than 1.