Concepts / Returns and Value-Function Updates

Returns and Value-Function Updates

λ identifies important endpoint cases of the λ-return.

  • Programming

Why the Endpoints Matter

The λ-return has two especially important settings: λ = 1 and λ = 0. These settings provide the clearest way to understand how the λ-return determines a value-function update. The essential task is to simplify the λ-return for the chosen value of λ and then identify the return that remains.

Classify a case from the return produced after the λ-return is simplified, not from the value of λ alone.

Following λ Through the Update

λ is useful because it identifies important endpoint cases of the λ-return. Start with a selected λ value, simplify the λ-return, and inspect the resulting return. If the result is the conventional return, the associated algorithm is Monte Carlo. If the result is the one-step return, the associated method is one-step TD. This classification depends on the return produced by the simplification.

selectsimplifyidentifyλ settingλ = 1 or λ = 0λ-returnsimplifyresulting returnconventional or one-stepupdate methodMonte Carlo or one-step TD
How does the selected λ value lead to the return and method used for the value-function update?

The λ = 1 Endpoint

When λ = 1, the λ-return simplifies to the conventional return. Backing up according to this return is identified as the Monte Carlo algorithm. This is a direct classification: λ = 1 produces the conventional return, and the conventional return leads to the Monte Carlo method.

simplifies tobacks up according toλ = 1conventional returnλ-return resultMonte Carlovalue-function update
What does the λ-return become when λ = 1, and which update method follows?

The λ = 1 result is stronger than saying the λ-return merely resembles Monte Carlo learning. After simplification, the return is the conventional return, and backing up according to it is identified as a Monte Carlo algorithm.

The λ = 0 Endpoint

When λ = 0, the relevant result is the one-step return, written in the source as G (1) t. That return leads to the one-step TD method. Therefore, the λ = 0 case must be classified as one-step TD, not as Monte Carlo.

simplifies toleads toλ = 0one-step returnG (1) tone-step TDvalue-function update
What does the λ-return become when λ = 0, and which update method follows?

Endpoint Classification Worked Example

Classifying Two λ Settings

Determine the return and update method for λ = 1 and λ = 0.

Classify λ = 1: The λ-return becomes the conventional return. The associated value-function update is therefore identified as the Monte Carlo algorithm.

Classify λ = 0: The λ-return becomes the one-step return, G (1) t. The associated value-function update is therefore identified as the one-step TD method.

Compare the results: The two settings produce different returns and therefore different associated methods. The λ = 1 case is not interchangeable with the λ = 0 case.

λ = 1 gives the conventional return and Monte Carlo; λ = 0 gives the one-step return and one-step TD.

λ settingλ-return becomesAssociated method
λ = 1Conventional returnMonte Carlo algorithm
λ = 0One-step return, G (1) tOne-step TD method
associated methodassociated methodλ = 1conventional returnλ = 0one-step returnMonte Carloone-step TD
What is the difference between the λ = 1 and λ = 0 endpoints?

Common Classification Mistakes

  • Calling λ = 1 one-step TD

    At λ = 1, the λ-return becomes the conventional return, which is associated with Monte Carlo.

    Fix: For λ = 1, identify the conventional return and then select Monte Carlo.

  • Calling λ = 0 Monte Carlo

    At λ = 0, the relevant result is the one-step return G (1) t, which leads to one-step TD.

    Fix: For λ = 0, identify the one-step return and then select one-step TD.

  • Choosing the method from λ without checking the return

    The classification rule is based on the return produced after the λ-return is simplified.

    Fix: Use the sequence: select λ, simplify the λ-return, identify the resulting return, then name the method.

Practice the Two Endpoints

EASY

For each setting, name the return produced by the λ-return and the associated value-function update method: λ = 1; λ = 0.

Hints
  • First identify whether the setting is the conventional-return endpoint or the one-step-return endpoint.
  • Then match the conventional return with Monte Carlo and the one-step return with one-step TD.

What do you think happens?

Before checking the answer, predict which method is associated with λ = 0.

  • Monte Carlo
  • One-step TD
Reveal answer

Answer: One-step TD

When λ = 0, the λ-return becomes the one-step return G (1) t, and that result leads to the one-step TD method.

Endpoint Summary

  1. λ identifies important endpoint cases of the λ-return.
  2. At λ = 1, the λ-return becomes the conventional return and the update is identified as Monte Carlo.
  3. At λ = 0, the λ-return becomes the one-step return G (1) t and the update is identified as one-step TD.
  4. The safest classification process is to simplify the λ-return first, identify the resulting return, and only then name the associated method.
  5. Monte Carlo at λ = 1 and one-step TD at λ = 0 are distinct endpoint cases and should not be confused.

Key Takeaways

  • λ is used to identify important endpoint cases of the λ-return.
  • λ = 1 produces the conventional return and therefore the Monte Carlo algorithm.
  • λ = 0 produces the one-step return G (1) t and therefore the one-step TD method.
  • Always classify the method by first identifying the return produced after simplifying the λ-return.