Returns and Value-Function Updates
λ identifies important endpoint cases of the λ-return.
Why the Endpoints Matter
The λ-return has two especially important settings: λ = 1 and λ = 0. These settings provide the clearest way to understand how the λ-return determines a value-function update. The essential task is to simplify the λ-return for the chosen value of λ and then identify the return that remains.
Classify a case from the return produced after the λ-return is simplified, not from the value of λ alone.
Following λ Through the Update
λ is useful because it identifies important endpoint cases of the λ-return. Start with a selected λ value, simplify the λ-return, and inspect the resulting return. If the result is the conventional return, the associated algorithm is Monte Carlo. If the result is the one-step return, the associated method is one-step TD. This classification depends on the return produced by the simplification.
The λ = 1 Endpoint
When λ = 1, the λ-return simplifies to the conventional return. Backing up according to this return is identified as the Monte Carlo algorithm. This is a direct classification: λ = 1 produces the conventional return, and the conventional return leads to the Monte Carlo method.
The λ = 1 result is stronger than saying the λ-return merely resembles Monte Carlo learning. After simplification, the return is the conventional return, and backing up according to it is identified as a Monte Carlo algorithm.
The λ = 0 Endpoint
When λ = 0, the relevant result is the one-step return, written in the source as G (1) t. That return leads to the one-step TD method. Therefore, the λ = 0 case must be classified as one-step TD, not as Monte Carlo.
Endpoint Classification Worked Example
Classifying Two λ Settings
Determine the return and update method for λ = 1 and λ = 0.
Classify λ = 1: The λ-return becomes the conventional return. The associated value-function update is therefore identified as the Monte Carlo algorithm.
Classify λ = 0: The λ-return becomes the one-step return, G (1) t. The associated value-function update is therefore identified as the one-step TD method.
Compare the results: The two settings produce different returns and therefore different associated methods. The λ = 1 case is not interchangeable with the λ = 0 case.
λ = 1 gives the conventional return and Monte Carlo; λ = 0 gives the one-step return and one-step TD.
| λ setting | λ-return becomes | Associated method |
|---|---|---|
| λ = 1 | Conventional return | Monte Carlo algorithm |
| λ = 0 | One-step return, G (1) t | One-step TD method |
Common Classification Mistakes
Calling λ = 1 one-step TD
At λ = 1, the λ-return becomes the conventional return, which is associated with Monte Carlo.
Fix:
For λ = 1, identify the conventional return and then select Monte Carlo.Calling λ = 0 Monte Carlo
At λ = 0, the relevant result is the one-step return G (1) t, which leads to one-step TD.
Fix:
For λ = 0, identify the one-step return and then select one-step TD.Choosing the method from λ without checking the return
The classification rule is based on the return produced after the λ-return is simplified.
Fix:
Use the sequence: select λ, simplify the λ-return, identify the resulting return, then name the method.
Practice the Two Endpoints
For each setting, name the return produced by the λ-return and the associated value-function update method: λ = 1; λ = 0.
Hints
- First identify whether the setting is the conventional-return endpoint or the one-step-return endpoint.
- Then match the conventional return with Monte Carlo and the one-step return with one-step TD.
What do you think happens?
Before checking the answer, predict which method is associated with λ = 0.
Reveal answer
Answer: One-step TD
When λ = 0, the λ-return becomes the one-step return G (1) t, and that result leads to the one-step TD method.
Endpoint Summary
- λ identifies important endpoint cases of the λ-return.
- At λ = 1, the λ-return becomes the conventional return and the update is identified as Monte Carlo.
- At λ = 0, the λ-return becomes the one-step return G (1) t and the update is identified as one-step TD.
- The safest classification process is to simplify the λ-return first, identify the resulting return, and only then name the associated method.
- Monte Carlo at λ = 1 and one-step TD at λ = 0 are distinct endpoint cases and should not be confused.
Key Takeaways
- λ is used to identify important endpoint cases of the λ-return.
- λ = 1 produces the conventional return and therefore the Monte Carlo algorithm.
- λ = 0 produces the one-step return G (1) t and therefore the one-step TD method.
- Always classify the method by first identifying the return produced after simplifying the λ-return.