Concepts / Greedy and Exploratory Policies

Greedy and Exploratory Policies

Off-policy learning separates the policy being evaluated from the policy generating experience.

  • Programming

Two Roles for Two Policies

Off-policy learning separates the policy whose value function is being learned from the policy that generates experience. The target policy is written as π. The followed policy is written as µ. This separation allows learning to use data produced while following µ while evaluating the value function associated with π.

generatesused to learndefinesπtarget policyexperiencegenerated datavalue functionassociated with πµfollowed policy
Which policy is being evaluated, which policy generates experience, and how does an episode move between the two?

The two policies are not interchangeable. µ supplies the experience, while π defines the value function of interest.

A Single Decision Point

Consider one state in which the target policy π selects one action, while the followed policy µ may select among several actions. The episode is generated according to µ, so the observed action comes from µ. However, the learning objective concerns π. The update therefore has to account for how the observed action relates to both policies.

selectsmay selectmay selectmay selectstateunder πaction Aselected by πstateunder µaction Apossible under µaction Bpossible under µaction Cpossible under µ
At the same state, how can a target policy choose one action while a followed policy chooses among several actions?

The important distinction is not simply that one policy selects one action and the other selects several. The important distinction is that the observed experience was generated by µ, whereas the value being learned belongs to π. Their difference must be represented in the update.

Why the Mismatch Matters

An action in the experience has a probability under the followed policy µ and a probability under the target policy π. Off-policy learning must account for the difference between those probabilities because the data came from µ but is being used to learn about π. Ignoring that difference would treat followed-policy experience as though it directly represented the target policy.

compare withcompare withcontributescontributesobserved actionfrom experienceπ probabilitytarget policypolicy correctionused in the updateµ probabilityfollowed policy
How does an observed action's probability under the target policy compare with its probability under the followed policy?

Following an N-Step Return

An n-step return is constructed from a sequence of n actions. Therefore, the off-policy correction focuses on the same n actions that were actually used to construct that return. The correction does not need to focus on unrelated actions outside this n-step sequence.

followed byfollowed bycontinues toreachescontributescontributescontributestime tstarting pointn-step returnconstructed from thesequenceaction 1first actionaction 2second actionaction nnth actiontime t+nend of sequence
Which actions were taken between time t and time t+n, and how do those actions contribute to the return being estimated?

Tracking the Relevant Actions

An experience sequence contains n actions between time t and time t+n. The sequence was generated by µ, but the value function being learned is associated with π. Which actions must the policy correction consider?

Identify the return: The return is an n-step return, so it was constructed from the n actions in the sequence.

Identify the source: The sequence came from the followed policy µ.

Identify the target: The value function of interest is associated with the target policy π.

Limit the correction to the sequence: Because these n actions constructed the return, the policy correction focuses on their relative probabilities under π and µ.

The relevant correction is tied to the n actions used to construct the n-step return.

Building ρt+n

The weighting term ρt+n combines the target-to-followed policy probability comparisons for the n actions in the return sequence. In other words, each of those n actions contributes a relative probability comparison between π and µ, and the combined weighting represents the policy mismatch for the whole n-action segment.

combinecombinecombineaction 1 ratioπ compared with µρt+ncombined weightingaction 2 ratioπ compared with µaction n ratioπ compared with µ
How are the target-to-followed policy probability ratios for the n actions combined into ρt+n?

The subscript t+n identifies the n-step segment associated with the weighting. It signals that the correction is connected to the actions between time t and time t+n.

From TD to Off-Policy TD

The ordinary n-step TD idea uses the n-step return to correct a value estimate. In the off-policy version, the weighting term ρt+n enters that correction. The n-step return still comes from the observed sequence, but its contribution to learning is adjusted to reflect the difference between the policy that generated the sequence and the policy whose value is being learned.

constructsdetermines relevant comparisonssupplies returnweights correctionn-action sequencegenerated by µn-step returnfrom the sequenceoff-policy TD updatevalue associated with πρt+npolicy correction
Where does ρt+n enter the n-step TD update, and how does it change the value correction?

Interpreting the Weighted Update

An n-step return is built from experience generated by µ, while the learner is estimating the value function associated with π. What changes when ρt+n is included?

Start with the observed return: The learner uses the n-step return constructed from the actions that occurred in the experience.

Compare the policies: For those same n actions, the learner considers their relative probabilities under the target policy π and the followed policy µ.

Combine the comparisons: The combined weighting is represented by ρt+n.

Apply the correction: The weighted return contributes to an off-policy n-step TD update for the value function associated with π.

Including ρt+n turns the n-step TD correction into an off-policy version that accounts for how the n observed actions relate to π and µ.

Common Reasoning Errors

  • Treating π and µ as the same policy.

    Off-policy learning specifically separates the target policy from the followed policy.

    Fix: Label π as the target policy and µ as the followed policy before tracing the update.

  • Replacing µ with π in the observed experience.

    The experience remains data produced by µ.

    Fix: Keep µ as the source of experience and use the correction to account for the difference from π.

  • Correcting for actions outside the n-step sequence.

    The n-step return uses a sequence of n actions, so the correction focuses on those same n actions.

    Fix: Trace the actions between time t and time t+n and connect the weighting to that segment.

  • Using ρt+n as if it were unrelated to the TD update.

    The weighting term ρt+n is what turns the n-step TD update into an off-policy version.

    Fix: Show how ρt+n weights the value correction based on the policy mismatch.

Practice Trace

MEDIUM

Suppose an episode is generated by µ, and a learner is estimating the value function associated with π. An n-step return uses the actions observed between time t and time t+n. Explain which policy supplies the data, which policy defines the value function, which actions the correction considers, and what role ρt+n plays.

Hints
  • Start by assigning the roles of π and µ.
  • Limit the action comparison to the n actions used to construct the return.
  • Describe ρt+n as the combined policy correction for that action sequence.

What do you think happens?

Before reading the explanation, predict whether an off-policy n-step correction should focus on every action in an entire episode or on the n actions used to construct the particular return.

  • Every action in the entire episode
  • The n actions used to construct the return
  • Only the first action, regardless of n
Reveal answer

Answer: The n actions used to construct the return

An n-step return uses a sequence of n actions, so the policy correction focuses on those same n actions. Their relative probabilities under π and µ are combined in ρt+n.

Key Takeaways

  1. Off-policy learning separates the target policy π from the followed policy µ.
  2. µ generates the experience, while π defines the value function being learned.
  3. Because the data comes from µ but the objective concerns π, the update must account for the difference between the policies.
  4. An n-step return uses n actions, so the correction focuses on those same n actions.
  5. The weighting term ρt+n combines the relevant policy comparisons and turns the n-step TD update into an off-policy version.

Key Takeaways

  • π is the target policy being evaluated, and µ is the followed policy that generates experience.
  • Off-policy updates must account for the mismatch between the policy producing the data and the policy whose value function is being learned.
  • The correction for an n-step return concerns the n actions used to construct that return.
  • ρt+n combines the relevant target-to-followed policy probability comparisons and weights the n-step TD correction.