Off-policy Learning with n-step Tree Backup
Expected action value is defined as a scalar random variable under the target policy.
Why the Formulation Starts with Two Quantities
The n-step tree-backup formulation is easier to express when it is built in stages. First, an expected action value is introduced under the target policy. Next, a form of TD error is introduced as a separate scalar random variable. The n-step tree-backup return is then defined using both quantities.
The order matters: expected action value first, TD error second, and n-step return third. Neither of the first two quantities is itself the final return.
The Policy-Based Quantity
The first ingredient is an expected action value associated with the target policy. The formulation treats this quantity as a scalar random variable. This treatment is not presented as the complete tree-backup return. It is a setup step that gives the later definition a compact quantity to use.
Expected action value: a scalar random variable defined under the target policy and used as one ingredient in the later definition of the n-step tree-backup return.
Backing Up through Policy Branches
The name tree backup highlights that the formulation is connected to the target policy's action possibilities. At each point in the conceptual tree, the target policy provides the policy context for the expected action value. The important lesson here is not a missing numerical equation; it is the construction pattern: policy-based expected quantities are backed up through multiple steps and are combined with TD-error quantities when the n-step return is defined.
The branches in this mental model are not a replacement for the formal definition. They illustrate why a target-policy expected action value is introduced before the multi-step return is defined.
The TD-Error Ingredient
The second ingredient is a form of TD error introduced as a separate scalar random variable. Its role is to provide a correction quantity for the tree-backup construction. The source distinguishes this TD-error quantity from the expected action value and then uses both quantities in defining the n-step return.
TD error in this formulation: a separately defined scalar random variable that is used together with expected action value when defining the n-step tree-backup return.
A Symbolic Construction Walkthrough
Building the Return in the Correct Order
Trace the conceptual construction of an n-step tree-backup return without treating any ingredient as the final result.
1. Identify the policy-based quantity: Begin with the expected action value. Record that it is associated with the target policy and is being treated as a scalar random variable.
2. Introduce the correction quantity: Define a form of TD error as a separate scalar random variable. Do not rename the expected action value or substitute one quantity for the other.
3. Extend across multiple steps: Use the policy-based quantity and the TD-error quantity as ingredients in the multi-step tree-backup construction.
4. Name the result correctly: Only the quantity defined after combining the two ingredients in the tree-backup formulation is the n-step return.
The construction has three distinct levels: expected action value, TD error, and the final n-step tree-backup return.
This walkthrough is intentionally symbolic. The supplied source establishes the construction order and the roles of the quantities, but it does not reproduce the actual equations. Therefore, the reliable lesson is how to classify each object and when it enters the construction, not a fabricated numerical calculation.
Common Category Errors
Treating the expected action value as the final n-step return.
It is a setup quantity and one ingredient in the later definition, not the complete return.
Fix:
Reserve the term n-step return for the quantity defined after the expected action value and TD error are used together.Treating TD error and expected action value as the same object.
The formulation introduces them as separate scalar random variables.
Fix:
Track them separately: expected action value supplies the target-policy quantity, while TD error supplies a distinct correction quantity.Describing the expected action value without mentioning the target policy.
The source specifically associates it with the target policy.
Fix:
State that it is an expected action value under the target policy.Jumping directly to the final return and skipping the two-quantity setup.
The compact tree-backup formulation depends on introducing those two scalar random variables first.
Fix:
Follow the construction order: expected action value, TD error, then n-step return.
Practice the Construction Order
Write a three-line explanation of the tree-backup setup. Line 1 should identify the expected action value and its policy association. Line 2 should identify the TD error as a separate scalar random variable. Line 3 should explain what is defined using both quantities.
Hints
- Include the phrase under the target policy for the expected action value.
- Do not describe the TD error as a second name for the expected action value.
- The final line should name the n-step tree-backup return.
What do you think happens?
Which item is the final n-step tree-backup return: the expected action value, the TD error, or the quantity defined using both?
Reveal answer
Answer: The quantity defined using both
The expected action value and TD error are separate scalar random variables introduced as ingredients. The source states that the n-step return is then defined using both quantities.
The Three-Level Mental Model
- Expected action value is introduced as a scalar random variable under the target policy.
- A form of TD error is introduced separately as another scalar random variable.
- The n-step tree-backup return is defined using both quantities.
- Neither expected action value nor TD error should be confused with the final n-step return.
- The safest reading order is policy-based quantity, TD-error quantity, and then multi-step return.
Key Takeaways
- Expected action value is a target-policy action-value quantity represented as a scalar random variable.
- TD error is a separate scalar random variable introduced for the tree-backup formulation.
- The n-step return is constructed from both quantities rather than being identical to either one.
- Keeping the three levels separate prevents the main category error in this formulation.