Concepts / n-step Action-Value Backups with Q(σ)

n-step Action-Value Backups with Q(σ)

ε-greedy behavior is defined relative to the current Q-values of the target policy.

  • Programming

Two Policies in One Learning Process

Off-policy n-step Q(σ) separates two jobs that are easy to confuse. The behavior policy μ selects the actions that generate experience. The target policy π determines the policy whose action values are being learned. The algorithm therefore collects a trajectory with one policy and uses that trajectory to update estimates for another policy.

selectsdetermines values learnedμexperience collectiontrajectory actionsobserved experienceπlearned policyQ-valuesaction values
Which policy generates the trajectory, and which policy determines the action values being learned?

Defining the Target from Q

The target policy π is initialized as ε-greedy with respect to Q, unless a fixed target policy is supplied. This means that the current Q-values provide the reference for defining the target policy's action preferences. As Q changes, an ε-greedy target policy defined from Q can change as well.

ε-greedy with respect to Q: a target-policy choice rule whose action preferences are defined relative to the current Q-values.

For a particular state, first inspect the available Q-values. Those values establish which actions the target policy treats as preferred and which actions are non-preferred. The source material establishes this relationship between π and Q, but it does not specify a numerical ε value or an exact probability-allocation formula. Therefore, an implementation must use the ε-greedy convention chosen for that implementation rather than assuming a probability from the name alone.

Q referenceQ referenceQ referenceaction a1Q(s,a1)πε-greedy preferencesaction a2Q(s,a2)action a3Q(s,a3)
Given current Q-values, which actions are treated as preferred when the target policy π is defined?

Walking Through the Backup

An n-step Q(σ) backup follows experience across multiple time steps before changing the current action-value estimate. At each backup stage, Q(σ) combines two possible ways of continuing: it can use a sampled reward-and-action continuation from the trajectory, or it can use an expected action-value continuation under the target policy. The parameter σ controls this mixture between sampled and expected continuation.

The important off-policy point is that the trajectory was generated by μ even though the expected action-value continuation is associated with π. The update therefore cannot treat the observed behavior-policy path as automatically representative of the target-policy path. Importance sampling supplies the connection between them.

trace experiencecontinuesampleexpectcontinuecontinuebackupcurrent Q(s,a)estimate to updatetime step 1observed reward and actionσ mixturesample or expectationsampled continuationtrajectory under μtime step nbackup endpointQ-value updaterevised estimateexpected continuationaction values under π
How does the backup move through several time steps, and where does Q(σ) choose sampled or expected continuation?

Connecting μ to π

Importance sampling connects the behavior-policy experience to the target-policy update. The algorithm uses ρ to correct the update when actions came from μ but the learning objective is defined by π. The correction is based on the relationship between the policies' action probabilities, and the relevant correction can accumulate across the steps included in the n-step backup.

This makes the action-probability trace a central debugging path. For every action used by the backup, identify the probability assigned by μ and the probability assigned by π. Then check how the corresponding ρ values are accumulated over the backup. A mismatch in either policy probability or the set of steps included in that accumulation can change the resulting update.

selected underevaluated underpolicy relationshippolicy relationshipacross stepscorrectsobserved actiontrajectory dataμ probabilitybehavior likelihoodρpolicy correctionratio accumulationbackup stepsQ-value updatetarget-policy learningπ probabilitytarget likelihood
How do action-probability relationships connect a trajectory generated by μ to an update for π?

Trace one action at a time. Record the action selected by μ, look up its probability under μ, look up its probability under π, and then verify that the corresponding ρ contribution is included at the intended backup step.

Tracing a Policy Mismatch

A generated mismatch trace

Suppose a trajectory contains an action selected by the behavior policy μ. The target policy π assigns that action a different probability. Explain why the Q(σ) update may differ from an update that ignored the policy distinction.

1. Identify the source of the action: The action belongs to experience generated under μ. It should not be described as though π selected it.

2. Inspect both policy probabilities: Compare the probability that μ assigns to the observed action with the probability that π assigns to the same action.

3. Check the correction: Because the learning target belongs to π, the importance-sampling correction ρ must connect the behavior-policy action to the target-policy update.

4. Check the backup range: For an n-step update, verify which actions contribute to the accumulated correction. A mismatch at an included step can change the update; a step outside the correction range should not be silently included.

The update can differ from the expected result when the implementation uses the wrong policy probability, omits a required ρ contribution, or accumulates the correction over the wrong backup steps.

selectscomparecontributes probabilityidentifies correctionincluded at stepchangesμselects actionobserved actiontrajectory stepπtarget probabilityρ contributionpolicy correctionn-step backupaccumulated correctionQ-value updatepossibly different result
At which action and backup step can a difference between μ and π alter the update?

Debugging Policy Probabilities

  • Treating μ and π as the same policy

    Off-policy n-step Q(σ) explicitly separates experience collection under μ from learning under π.

    Fix: Label every action by its generating policy and use π when reasoning about the action values being learned.

  • Defining π without looking at Q

    The target policy is ε-greedy with respect to the current Q-values.

    Fix: Inspect the current Q-values before determining the target policy's action preferences.

  • Ignoring ρ after observing a policy mismatch

    Importance sampling connects the two policies by correcting the update with ρ.

    Fix: Trace the policy probabilities and verify the accumulated ρ contribution for the n-step backup.

  • Checking only the first action

    The correction can accumulate across the steps included in the n-step backup.

    Fix: Audit every action and every backup step that contributes to the correction.

  • Allowing μ to assign zero probability to an encountered state-action possibility

    The algorithm requires μ to assign positive probability to every state-action pair.

    Fix: Check the behavior-policy support condition before trusting the off-policy update.

Practice Trace

MEDIUM

A trajectory was generated by μ, but the update is intended to learn Q-values for π. Write a four-line audit: identify the action-generating policy, identify the target policy, list the policy probabilities that must be compared, and state where the accumulated ρ correction must be checked.

Hints
  • Start with the roles of μ and π, not with the Q-value update.
  • Use the observed action as the common object whose probabilities are compared.
  • Include every backup step that contributes to the n-step correction.
  1. A reliable audit follows the data in order: action selection under μ, current Q-values defining ε-greedy preferences for π, policy-probability comparison, ρ accumulation, and finally the Q-value update.

Key Takeaways

  • μ generates the trajectory, while π identifies the policy whose action values are learned.
  • An ε-greedy target policy is defined relative to the current Q-values.
  • Q(σ) combines sampled trajectory information with expected action-value continuation under the target policy.
  • Importance sampling uses ρ to connect behavior-policy experience to the target-policy update.
  • A mismatch can change the update when policy probabilities or the steps included in ratio accumulation are handled incorrectly.