Concepts / Off-policy Learning with n-step Tree Backup

Off-policy Learning with n-step Tree Backup

Expected action value is defined as a scalar random variable under the target policy.

  • Programming

Why the Formulation Starts with Two Quantities

The n-step tree-backup formulation is easier to express when it is built in stages. First, an expected action value is introduced under the target policy. Next, a form of TD error is introduced as a separate scalar random variable. The n-step tree-backup return is then defined using both quantities.

The order matters: expected action value first, TD error second, and n-step return third. Neither of the first two quantities is itself the final return.

used in definitionused in definitionExpected action valueunder target policyTD errorseparate scalarn-step returntree-backup result
How are the two scalar random variables combined conceptually to produce the final n-step tree-backup return?

The Policy-Based Quantity

The first ingredient is an expected action value associated with the target policy. The formulation treats this quantity as a scalar random variable. This treatment is not presented as the complete tree-backup return. It is a setup step that gives the later definition a compact quantity to use.

Expected action value: a scalar random variable defined under the target policy and used as one ingredient in the later definition of the n-step tree-backup return.

defines contextserves as ingredientTarget policypolicy contextExpected action valuescalar random variablen-step returndefinitionlater use
What does the scalar expected action value represent, and where does its policy association come from?

Backing Up through Policy Branches

The name tree backup highlights that the formulation is connected to the target policy's action possibilities. At each point in the conceptual tree, the target policy provides the policy context for the expected action value. The important lesson here is not a missing numerical equation; it is the construction pattern: policy-based expected quantities are backed up through multiple steps and are combined with TD-error quantities when the n-step return is defined.

target-policy branchtarget-policy branchbacked upbacked upCurrent decisiontarget-policy contextAction branch Aexpected contributionNext backup levelcontinued constructionAction branch Bexpected contribution
How does the target policy provide the branching context for expected contributions in a tree-backup formulation?

The branches in this mental model are not a replacement for the formal definition. They illustrate why a target-policy expected action value is introduced before the multi-step return is defined.

The TD-Error Ingredient

The second ingredient is a form of TD error introduced as a separate scalar random variable. Its role is to provide a correction quantity for the tree-backup construction. The source distinguishes this TD-error quantity from the expected action value and then uses both quantities in defining the n-step return.

TD error in this formulation: a separately defined scalar random variable that is used together with expected action value when defining the n-step tree-backup return.

ingredientingredientExpected action valuepolicy-based ingredientExpected action valuepolicy-based ingredientTD errorseparate scalarn-step returndefined using both
What changes in the construction when a TD error is introduced at each step?

A Symbolic Construction Walkthrough

Building the Return in the Correct Order

Trace the conceptual construction of an n-step tree-backup return without treating any ingredient as the final result.

1. Identify the policy-based quantity: Begin with the expected action value. Record that it is associated with the target policy and is being treated as a scalar random variable.

2. Introduce the correction quantity: Define a form of TD error as a separate scalar random variable. Do not rename the expected action value or substitute one quantity for the other.

3. Extend across multiple steps: Use the policy-based quantity and the TD-error quantity as ingredients in the multi-step tree-backup construction.

4. Name the result correctly: Only the quantity defined after combining the two ingredients in the tree-backup formulation is the n-step return.

The construction has three distinct levels: expected action value, TD error, and the final n-step tree-backup return.

This walkthrough is intentionally symbolic. The supplied source establishes the construction order and the roles of the quantities, but it does not reproduce the actual equations. Therefore, the reliable lesson is how to classify each object and when it enters the construction, not a fabricated numerical calculation.

used to defineused to defineExpected action valueunder target policyTD errorseparate scalarn-step returntree-backup result
How do expected action value and TD error differ from the final n-step tree-backup return?

Common Category Errors

  • Treating the expected action value as the final n-step return.

    It is a setup quantity and one ingredient in the later definition, not the complete return.

    Fix: Reserve the term n-step return for the quantity defined after the expected action value and TD error are used together.

  • Treating TD error and expected action value as the same object.

    The formulation introduces them as separate scalar random variables.

    Fix: Track them separately: expected action value supplies the target-policy quantity, while TD error supplies a distinct correction quantity.

  • Describing the expected action value without mentioning the target policy.

    The source specifically associates it with the target policy.

    Fix: State that it is an expected action value under the target policy.

  • Jumping directly to the final return and skipping the two-quantity setup.

    The compact tree-backup formulation depends on introducing those two scalar random variables first.

    Fix: Follow the construction order: expected action value, TD error, then n-step return.

Practice the Construction Order

EASY

Write a three-line explanation of the tree-backup setup. Line 1 should identify the expected action value and its policy association. Line 2 should identify the TD error as a separate scalar random variable. Line 3 should explain what is defined using both quantities.

Hints
  • Include the phrase under the target policy for the expected action value.
  • Do not describe the TD error as a second name for the expected action value.
  • The final line should name the n-step tree-backup return.

What do you think happens?

Which item is the final n-step tree-backup return: the expected action value, the TD error, or the quantity defined using both?

  • The expected action value
  • The TD error
  • The quantity defined using both
Reveal answer

Answer: The quantity defined using both

The expected action value and TD error are separate scalar random variables introduced as ingredients. The source states that the n-step return is then defined using both quantities.

The Three-Level Mental Model

  1. Expected action value is introduced as a scalar random variable under the target policy.
  2. A form of TD error is introduced separately as another scalar random variable.
  3. The n-step tree-backup return is defined using both quantities.
  4. Neither expected action value nor TD error should be confused with the final n-step return.
  5. The safest reading order is policy-based quantity, TD-error quantity, and then multi-step return.

Key Takeaways

  • Expected action value is a target-policy action-value quantity represented as a scalar random variable.
  • TD error is a separate scalar random variable introduced for the tree-backup formulation.
  • The n-step return is constructed from both quantities rather than being identical to either one.
  • Keeping the three levels separate prevents the main category error in this formulation.