Concepts / Action-Value Methods in Reinforcement Learning

Action-Value Methods in Reinforcement Learning

n-step Q(σ) is a unifying framework for action-value backups.

  • Programming

One Backup, Several Choices

An n-step action-value backup looks ahead through several successive decision points to improve an estimate of how valuable an action is. Earlier methods made different choices about the actions encountered during that lookahead. n-step Sarsa follows sampled actions, tree backup accounts for possible actions through expectations, and n-step Expected Sarsa uses a particular expected-action pattern. n-step Q(σ) provides one framework for describing these choices.

The Q(σ) Backup Path

Suppose a backup examines several future decision points. At every point, it can either commit to the action that was sampled in the trajectory or replace that single action with an expectation over actions. Q(σ) makes this decision separately for each step. Therefore, one backup can contain a mixture: one future step may use a sampled action, while another may use an expected action value.

observecontinuemove forwardcontinueform targetCurrent actionvalue estimateRewardfirst stepRewardnext stepBootstrap valuen-step targetSample or expectσtSample or expectσt+1
How does an n-step backup combine rewards, sampled actions, and expected action values across multiple future steps?

The subscript in σt matters. It indicates that the degree of sampling can change from one time step to another. The choice may also be made as a function of the state, the action, or the state-action pair at that time.

Reading the Sampling Parameter

σt describes the degree of sampling used at step t. When σt = 1, the backup uses full sampling: it follows the sampled action. When σt = 0, it uses pure expectation: it accounts for actions through an expected action value. Values between 0 and 1 represent an intermediate degree of sampling.

less samplingmore expectationσt = 1sampled action0 < σt < 1mixed samplingσt = 0expected action value
What changes in the backup at step t as σt moves from fully sampled action selection to fully expected action values?

A mixed three-step backup

Consider a three-step backup whose sampling choices are σt = 1, σt+1 = 0, and σt+2 = 1.

First step: Because σt = 1, this step uses the action sampled in the trajectory.

Second step: Because σt+1 = 0, this step uses an expectation over possible actions rather than committing to one sampled action.

Third step: Because σt+2 = 1, this step returns to full sampling.

The backup is neither fully sampled nor fully expected. It combines the two approaches at different points in the same lookahead.

Recovering the Earlier Backups

The earlier n-step action-value methods can be understood as particular patterns of σ choices. n-step Sarsa uses sampled actions throughout the lookahead, corresponding to full sampling at the relevant steps. Tree backup uses expectations throughout the lookahead, corresponding to pure expectation at those steps. n-step Expected Sarsa uses an established pattern that includes an expected action-value backup rather than relying on a sampled action at the expected-action part of the backup.

particular σ patternparticular σ patternparticular σ patternn-step Sarsasampled actionsn-step Q(σ)variable σ choicesTree backupexpectationsn-step ExpectedSarsaexpected-action pattern
Which choices of σ reproduce Sarsa, tree backup, or Expected Sarsa, and how do their backup paths differ?
MethodAction handling in the lookaheadRelationship to Q(σ)
n-step SarsaUses sampled actionsFull-sampling pattern
Tree backupUses expectations over actionsPure-expectation pattern
n-step Expected SarsaUses its established expected-action backup patternA particular Q(σ) pattern
n-step Q(σ)Can choose sampling or expectation separately at each stepUnifying framework

A Backup Trace

Imagine a trajectory with several successive decision points. At the first point, the backup follows the action that the agent actually sampled. At the second point, it accounts for possible actions through an expectation. At the third point, it again follows a sampled action. This trace illustrates the practical meaning of a step-dependent σ: the backup can switch its treatment of actions as it moves forward.

What do you think happens?

If σt changes from 1 to 0 at the next decision point, what changes in the backup?

  • The backup changes from using a sampled action to using an expected action value
  • The backup stops using rewards
  • The policy is permanently changed
  • The agent is evaluated over its entire lifetime
Reveal answer

Answer: The backup changes from using a sampled action to using an expected action value.

σt = 1 represents full sampling, while σt = 0 represents pure expectation. The change affects how the next action contribution is handled; it does not by itself describe a policy update or lifetime evaluation.

The rewards encountered during the lookahead remain part of the backup. The σ choice controls how action values are handled at each future decision point.

Evolution Without Value Estimates

Evolutionary methods take a different route through reinforcement learning. Many methods estimate how valuable states or actions are. Evolutionary methods instead use policy search and reward comparison. Each candidate agent follows its policy while interacting with the environment, receives reward for that lifetime of behavior, and is judged by the result. The strongest candidates become the basis for further search.

each candidate actsproduce outcomejudge resultsfavor stronger candidatescontinue searchCandidate policiesgroup of agentsLifetime interactionpolicy guides behaviorRewardcomplete behaviorCompare candidatesbetter reward favoredSelected policiesnext search stage
How do evolutionary methods generate agents, evaluate their lifetime behavior, compare fitness, and select agents for the next generation without estimating state or action values?

The important evaluation unit is the agent's complete lifetime behavior. The method asks how well an entire policy performed during its interaction period, rather than estimating the value attached to one particular state or state-action pair.

Lifetime Scores and Value Estimates

evaluate whole behaviorestimate local valueComplete policylifetime rewardBehavior scorecompare candidatesState-action pairvalue estimateAction valueparticular choice
What is the difference between scoring an agent's complete lifetime behavior and estimating the expected return from a particular state or state-action pair?
QuestionEvolutionary methodValue-function method
What is evaluated?An entire policy's behavior over its lifetimeA state or action estimate
What is compared?Rewards obtained by candidate agentsEstimated values associated with states or actions
Is a value function required?NoYes, for the value-estimation approach
How does search progress?Better candidates are favored for further searchValue estimates guide action-value improvement

Selection Across Generations

Consider a collection of non-learning agents. Each agent uses a different policy during its lifetime and does not improve through learning while it interacts with the environment. After the interaction period, the method compares the reward obtained by the complete behaviors. Candidates associated with better reward are selected as the basis for the next stage of policy search, while weaker candidates are not favored.

run policiesproduce scoresfavor better rewardcontinue policy searchGeneration 1candidate policiesEvaluatelifetime behaviorCompare rewardcandidate fitnessSelectstronger candidatesGeneration 2continued search
How do policies move through a population as agents are evaluated, compared, selected, and varied across generations?

The evolutionary analogy is specific: an individual agent can display successful behavior without learning during its own lifetime, while the broader population search favors policies associated with better results.

When Evolutionary Search Fits

Evolutionary methods may be effective when the policy space is small, when good policies are common or easy to find within that space, or when substantial time is available for searching. These conditions make it more plausible that repeated comparison and selection will encounter useful policies.

They can also be useful when the learning agent cannot accurately sense the state of its environment. Methods that depend on accurate environmental sensing may have difficulty in that situation, while evolutionary evaluation can still compare the behavior produced by different policies according to the reward obtained.

Mistakes in Interpretation

  • Treating σt as a policy or a reward value

    σt controls whether the backup uses a sampled action or an expectation at a particular step.

    Fix: Read σt as the degree of sampling: 1 means full sampling, 0 means pure expectation.

  • Assuming one sampling choice must be used at every step

    The subscript t allows the sampling degree to vary from one step to another.

    Fix: Inspect σt separately at each relevant step.

  • Confusing a lifetime behavior score with an action-value estimate

    Evolutionary methods judge complete policy behavior over an interaction period.

    Fix: Keep the evaluation units separate: whole-policy lifetime reward versus a value estimate for a state or action.

  • Assuming evolutionary agents learn during their own lifetimes

    The source describes candidates that follow their policies during their lifetimes and are compared afterward.

    Fix: Separate individual lifetime behavior from the broader population-level search.

Check Your Understanding

MEDIUM

A backup uses σt = 1 at its first future decision point, σt+1 = 0 at its second, and an intermediate value at its third. Describe how action choices are handled at all three points. Then explain whether this backup is best described as a fixed Sarsa-style backup, a fixed tree-backup-style backup, or a mixed Q(σ) backup.

Hints
  • Start with the meaning of σ = 1.
  • Then interpret σ = 0.
  • An intermediate value represents a degree between full sampling and pure expectation.
  • Because the choices differ across steps, focus on Q(σ)'s ability to vary the decision.
EASY

A researcher has a group of candidate policies but does not build state-value or action-value estimates. Each candidate interacts with the environment, receives a reward for its lifetime behavior, and is compared with the others. Identify the approach and explain what is being selected.

Hints
  • Look at whether the method estimates local values or compares complete behaviors.
  • The candidates are policies or agents in a population.
  • Selection favors candidates associated with more reward.

Key Takeaways

  • n-step Q(σ) unifies n-step action-value backups by deciding separately at each step whether to use a sampled action or an expected action value.
  • σt = 1 means full sampling, σt = 0 means pure expectation, and intermediate values represent a degree of sampling between those endpoints.
  • n-step Sarsa, tree backup, and n-step Expected Sarsa can be described as particular patterns of σ choices.
  • Evolutionary methods compare complete policy behaviors over agent lifetimes instead of estimating the value of individual states or actions.
  • Evolutionary search may be effective when the policy space is small or searchable, sufficient search time is available, or accurate environmental sensing is difficult.