Forward View of a Learning Algorithm
Learning is postponed until the episode has finished.
The Timing Question
The forward view is mainly a way to understand when learning happens and what information each update uses. In the off-line λ-return algorithm, the weight vector is not changed while the episode is still running. Learning is postponed until the episode has finished. The completed episode then supplies the future information needed to construct the λ-return targets.
During the Episode
The phrase off-line describes the algorithm's timing. While the episode is being experienced, the algorithm records the visited states and the future outcomes that will eventually be available. It does not yet change the weight vector. This matters because the target for a state depends on information later in the episode, and that information is not complete until the episode ends.
Once the episode has finished, the algorithm can process the visited states. For a selected state, it looks forward through the remainder of the episode. It uses the future rewards and future states to determine the λ-return for that state. It then moves to another visited state and performs the corresponding update. Later states can appear in the lookahead for several earlier states because they are future states relative to each of those earlier states.
Constructing the λ-Return
The λ-return is the target used by the off-line semi-gradient updates. It is formed from information found by looking forward from the state being processed. The important idea is not to treat the current state in isolation: future rewards and future states from the completed episode are combined to decide how that state's estimate should be adjusted.
Keep two roles separate: the λ-return specifies what value the update is trying to use as its target, while the off-line algorithm specifies when the update is carried out. The target comes from the forward-looking calculation; the weight-vector changes wait until the episode is complete.
Between Two Learning Views
| Learning approach | How much future information the target reflects | Relationship to the λ-return |
|---|---|---|
| One-step TD | A short, one-step lookahead | One end of the continuum |
| λ-return | A combination of information from future rewards and states | A position between the two familiar approaches |
| Monte Carlo | The completed episode's future outcome | The other end of the continuum |
The λ-return is useful because it occupies a continuum between one-step TD methods and Monte Carlo methods. Changing λ changes how the target balances information from shorter and longer lookaheads. A target closer to the one-step TD side relies more on a short lookahead, while a target closer to the Monte Carlo side reflects more of the completed episode. The forward view therefore gives a way to understand λ as controlling the character of the target, without changing the central off-line timing rule.
A State-by-State Walkthrough
Processing three visited states
Imagine an episode that visits state A, then state B, then state C, and finally finishes. Describe how the forward view processes these states after the episode ends.
Finish the episode: While the episode is unfolding, the off-line algorithm leaves the weight vector alone. It waits until all future rewards and states in the episode are available.
Process state A: Look forward from A through the remainder of the completed episode. The future rewards and states provide the information used to construct A's λ-return, which becomes the target for A's update.
Process state B: Move to B and perform the corresponding target-based update. The future part of the episode relative to B is shorter than it was relative to A.
Process state C: Process C in the same forward-view manner. A later state may have appeared in the lookahead for an earlier state, but the algorithm now focuses on the state currently being processed.
The episode supplies the information first, and the sequence of λ-return updates follows afterward. The weight vector is changed only during this post-episode update sequence.
Mistakes About the Forward View
Assuming the weight vector changes during the episode
The off-line algorithm leaves the weight vector alone while the episode is in progress.
Fix:
Separate data collection from learning: complete the episode first, then perform the sequence of updates.Treating the λ-return as the current value estimate
The λ-return is the target used by the semi-gradient update; it is calculated from future information in the completed episode.
Fix:
Keep the current estimate and the λ-return target conceptually separate.Looking only at the selected state
The forward view is defined by looking ahead through the rest of the episode.
Fix:
For each visited state, inspect the future part of the completed episode before identifying its target.Confusing the λ-return with only one of the endpoint methods
The λ-return provides a position between one-step TD and Monte Carlo methods.
Fix:
Think of λ as controlling the balance between shorter and longer future lookaheads.
Check Your Understanding
An episode has just started and two states have been visited. Has the off-line λ-return algorithm changed its weight vector yet? After the episode finishes, what information should be considered when constructing the target for the first visited state?
Hints
- Recall the difference between during-episode processing and post-episode learning.
- The first visited state is evaluated using the future part of the completed episode.
- Name the target used by the semi-gradient update.
Key Takeaways
- The off-line λ-return algorithm leaves its weight vector unchanged during the episode.
- After the episode finishes, the completed episode supplies the future rewards and states needed for learning.
- The forward view processes each visited state by looking ahead through the remainder of the episode.
- The λ-return is the target for the sequence of semi-gradient updates.
- Changing λ places the target along a continuum between one-step TD and Monte Carlo methods.