Value Estimation Across Episode Horizons
The h-truncated λ-return method expands its available information one horizon at a time.
One Episode, Several Passes
The h-truncated λ-return method does not process an episode once and wait until the end before making all updates. Instead, it expands its available information one horizon at a time. At horizon 1, it makes the updates possible from the earliest available data. At horizon 2, it revisits the episode with more data, revises the earlier target, and updates the next state as well. The same pattern continues as the horizon advances.
A later horizon is a new pass over the episode, not merely the next step appended to the previous pass.
What Each Horizon Knows
Before learning begins for a new episode, the algorithm starts with θ₀, the weight vector inherited from the end of the previous episode. When the data horizon reaches time step 1, the algorithm has observed R₁ and reached S₁. For the estimate at step 0, the available target is therefore the one-step return: it uses R₁ and then bootstraps from the estimate of S₁ generated with θ₀.
When the horizon advances, additional episode data allows targets for more earlier time steps. At horizon 2, the first update uses the improved target G λ|2 0 for the state at step 0, and the next update uses G λ|2 1 for the state at step 1. At horizon 3, the pass can use the targets for steps 0, 1, and 2. The essential change is not just the number of updates: the target available for an earlier state can also improve because the horizon contains more episode information.
Reading the Weight Notation
A single symbol such as θ₁ becomes ambiguous when the algorithm makes several passes. The notation θʰₜ resolves this ambiguity. The superscript h identifies the horizon of the pass, while the subscript t identifies the time position within that pass. Thus, θʰₜ records both where the vector occurs inside a pass and which horizon-specific pass is being considered.
Every horizon-specific sequence begins with the inherited θ₀. The vector at the end of the pass is written θʰ, meaning θʰₕ. It is the final vector produced after all updates available at horizon h have been performed.
Tracing a Complete Horizon Pass
From Horizon 2 to Horizon 3
Trace how the algorithm constructs the horizon-2 and horizon-3 results without treating horizon 3 as a continuation of horizon 2.
Start horizon 2: Begin the horizon-2 sequence at the inherited θ₀, not at the final vector from horizon 1.
Update step 0: Use the improved target G λ|2 0 for the state at step 0.
Update step 1: Use G λ|2 1 for the state at step 1. The two ordered updates produce θ².
Start horizon 3: Begin a new horizon-3 sequence again at the inherited θ₀.
Perform three updates: Use the targets for steps 0, 1, and 2 at horizon 3 in order. The sequence finishes at θ³.
Each horizon has its own sequence from θ₀ to its final vector. Horizon 3 is a new three-update pass, not one extra update appended to the horizon-2 sequence.
This trace illustrates the central bookkeeping rule. The algorithm revisits the beginning because later data can construct targets for more earlier states. It also restarts from θ₀ so that the horizon-specific pass is defined consistently. The final vector of a previous horizon is not used as the starting vector of the next horizon within the same episode.
Mistakes in Horizon Tracing
Treating horizon 2 as a continuation of horizon 1.
Each horizon-specific pass starts from the inherited θ₀ and performs the updates available at that horizon.
Fix:
Restart the horizon-2 sequence at θ₀, then perform its two updates in order.Using one symbol such as θ₁ without recording the horizon.
The same time position can occur in multiple horizon-specific passes.
Fix:
Use θʰₜ: h identifies the pass horizon and t identifies the time position within that pass.Assuming later horizons only add updates for later states.
Later horizons can form better targets for earlier states because additional episode data is available.
Fix:
Revisit earlier states and use the target associated with the new horizon.Stopping at the first update instead of completing the horizon pass.
θ³ denotes the final vector after all updates available at horizon 3 have been performed.
Fix:
Perform the updates for steps 0, 1, and 2 in order before identifying the result as θ³.
When tracing the method, write the horizon above the pass and the time position below each vector. This keeps the pass identity separate from the position within that pass.
Practice the Handoff
A complete episode reaches horizon T. Describe the starting vector, the final vector at that horizon, and the vector used to begin the next episode.
Hints
- Every horizon-specific sequence begins with the inherited θ₀.
- At the final horizon, the completed pass ends at θᵀ.
- The next episode uses the previous episode's final weights as its θ₀.
What do you think happens?
If the algorithm is about to begin horizon 3 after completing horizon 2, which vector starts the horizon-3 sequence?
Reveal answer
Answer: The inherited θ₀.
Horizon 3 is a new horizon-specific sequence. It starts from θ₀ and performs three updates in order before producing θ³.
Episode-to-Episode Learning
Within one episode, the method repeatedly expands its horizon, revisiting earlier states as more information becomes available. Each pass starts from the inherited θ₀, uses the targets available at its own horizon, and ends with a horizon-specific final vector. At the complete horizon T, the final weights are θᵀ. Those weights become the θ₀ inherited by the next episode, connecting repeated passes within an episode to learning across episodes.
- The h-truncated λ-return method revisits an episode as its available data horizon grows.
- Later horizons can improve targets for earlier states and can update additional states.
- The notation θʰₜ identifies both the horizon h and the position t within that pass.
- Every horizon-specific pass starts from θ₀ and ends with its own final vector θʰ.
- The complete-horizon result θᵀ becomes θ₀ for the next episode.
Key Takeaways
- The method expands available episode information one horizon at a time.
- A later horizon is a fresh pass from θ₀, not a continuation from the previous horizon's final vector.
- The superscript in θʰₜ identifies the horizon, while the subscript identifies the time position within that pass.
- At horizon T, θᵀ is the final weight vector and is carried into the next episode as its inherited θ₀.