Concepts / Forward View of a Learning Algorithm

Forward View of a Learning Algorithm

Learning is postponed until the episode has finished.

  • Programming

The Timing Question

The forward view is mainly a way to understand when learning happens and what information each update uses. In the off-line λ-return algorithm, the weight vector is not changed while the episode is still running. Learning is postponed until the episode has finished. The completed episode then supplies the future information needed to construct the λ-return targets.

finishessupplies informationbecomes update targetEpisodefuture rewards and statesCompleted episodeinformation availableλ-return targetsone for each visited stateWeight vectorsequence of updates
When does the off-line algorithm change its weight vector?

During the Episode

The phrase off-line describes the algorithm's timing. While the episode is being experienced, the algorithm records the visited states and the future outcomes that will eventually be available. It does not yet change the weight vector. This matters because the target for a state depends on information later in the episode, and that information is not complete until the episode ends.

Once the episode has finished, the algorithm can process the visited states. For a selected state, it looks forward through the remainder of the episode. It uses the future rewards and future states to determine the λ-return for that state. It then moves to another visited state and performs the corresponding update. Later states can appear in the lookahead for several earlier states because they are future states relative to each of those earlier states.

look aheadfollow episodecontinue looking aheadcontributes informationVisited statestate being processedFuture rewardlater outcomeSuccessor statelater stateFuture rewardanother later outcomeλ-returntarget for the state
Which future rewards and successor states are considered when constructing a state's learning target?

Constructing the λ-Return

The λ-return is the target used by the off-line semi-gradient updates. It is formed from information found by looking forward from the state being processed. The important idea is not to treat the current state in isolation: future rewards and future states from the completed episode are combined to decide how that state's estimate should be adjusted.

combinedcombinedsupplies informationShort lookaheadnear futureLonger lookaheadmore future outcomesCompleted episodeavailable futureinformationλ-returnlearning target
How are different amounts of future information combined into the target used for learning?

Keep two roles separate: the λ-return specifies what value the update is trying to use as its target, while the off-line algorithm specifies when the update is carried out. The target comes from the forward-looking calculation; the weight-vector changes wait until the episode is complete.

Between Two Learning Views

Learning approachHow much future information the target reflectsRelationship to the λ-return
One-step TDA short, one-step lookaheadOne end of the continuum
λ-returnA combination of information from future rewards and statesA position between the two familiar approaches
Monte CarloThe completed episode's future outcomeThe other end of the continuum

The λ-return is useful because it occupies a continuum between one-step TD methods and Monte Carlo methods. Changing λ changes how the target balances information from shorter and longer lookaheads. A target closer to the one-step TD side relies more on a short lookahead, while a target closer to the Monte Carlo side reflects more of the completed episode. The forward view therefore gives a way to understand λ as controlling the character of the target, without changing the central off-line timing rule.

more future informationlonger lookaheadOne-step TDshort lookaheadλ-returncombined lookaheadMonte Carlocompleted episode
How does changing λ move the target between a one-step TD estimate and a full-episode Monte Carlo return?

A State-by-State Walkthrough

Processing three visited states

Imagine an episode that visits state A, then state B, then state C, and finally finishes. Describe how the forward view processes these states after the episode ends.

Finish the episode: While the episode is unfolding, the off-line algorithm leaves the weight vector alone. It waits until all future rewards and states in the episode are available.

Process state A: Look forward from A through the remainder of the completed episode. The future rewards and states provide the information used to construct A's λ-return, which becomes the target for A's update.

Process state B: Move to B and perform the corresponding target-based update. The future part of the episode relative to B is shorter than it was relative to A.

Process state C: Process C in the same forward-view manner. A later state may have appeared in the lookahead for an earlier state, but the algorithm now focuses on the state currently being processed.

The episode supplies the information first, and the sequence of λ-return updates follows afterward. The weight vector is changed only during this post-episode update sequence.

is adjusted toward targetserves as targetCurrent estimatestate value before updateλ-returntarget from futureinformationUpdated estimateafter semi-gradient update
How does the target relate to the current estimate during an off-line update?

Mistakes About the Forward View

  • Assuming the weight vector changes during the episode

    The off-line algorithm leaves the weight vector alone while the episode is in progress.

    Fix: Separate data collection from learning: complete the episode first, then perform the sequence of updates.

  • Treating the λ-return as the current value estimate

    The λ-return is the target used by the semi-gradient update; it is calculated from future information in the completed episode.

    Fix: Keep the current estimate and the λ-return target conceptually separate.

  • Looking only at the selected state

    The forward view is defined by looking ahead through the rest of the episode.

    Fix: For each visited state, inspect the future part of the completed episode before identifying its target.

  • Confusing the λ-return with only one of the endpoint methods

    The λ-return provides a position between one-step TD and Monte Carlo methods.

    Fix: Think of λ as controlling the balance between shorter and longer future lookaheads.

Check Your Understanding

MEDIUM

An episode has just started and two states have been visited. Has the off-line λ-return algorithm changed its weight vector yet? After the episode finishes, what information should be considered when constructing the target for the first visited state?

Hints
  • Recall the difference between during-episode processing and post-episode learning.
  • The first visited state is evaluated using the future part of the completed episode.
  • Name the target used by the semi-gradient update.

Key Takeaways

  • The off-line λ-return algorithm leaves its weight vector unchanged during the episode.
  • After the episode finishes, the completed episode supplies the future rewards and states needed for learning.
  • The forward view processes each visited state by looking ahead through the remainder of the episode.
  • The λ-return is the target for the sequence of semi-gradient updates.
  • Changing λ places the target along a continuum between one-step TD and Monte Carlo methods.