Value Backups
Backward focusing directs planning from a changed value toward states that may depend on it.
A Changed Value as a Starting Point
Planning does not always need to recompute everything when new information arrives. If the estimated value of one state changes, planning can begin there and examine the states and actions that could be affected. This backward direction is called backward focusing.
Backward focusing directs planning from a changed value toward states that may depend on it.
The central question is not which states exist everywhere in the problem. It is which actions lead directly into the state whose value changed. Those incoming actions are the first useful backups. Once their predecessor states are updated, any resulting value changes can make still earlier predecessor states relevant.
The First Useful Backups
Immediately after a state changes in value, focus on the one-step backups for actions that lead directly into that state. These backups use the changed state as the value information at the end of the one-step transition. Actions unrelated to the changed state are not the immediate focus of this propagation step.
Following a Value Change Backward
Suppose the estimated value of state C changes. Determine the order in which planning focuses on related actions and states.
Locate the changed state: Begin with state C because its value is the new information.
Find direct incoming actions: Identify actions that lead directly into C. Their one-step backups are immediately useful because they can use C's changed value.
Inspect predecessor states: If updating those incoming actions changes the values of their predecessor states, those predecessor states become the next possible sources of backward propagation.
Ignore unrelated actions for this step: Actions that do not lead directly into C are not the immediate focus of this propagation step.
The planning focus moves backward from C to its direct incoming actions and then, when needed, to earlier predecessor states.
Propagation Beyond Goal States
Backward focusing is not restricted to a single explicit goal state. The changed information may be a newly discovered reward or another changed value. What matters is that a state value has changed and that earlier states may depend on it through actions leading forward to that state.
The Q(σ) Choice at Each Step
n-step Q(σ) is a unifying framework for action-value backups. At each step of an n-step backup, it chooses how much to rely on a sampled action and how much to use an expectation over possible actions.
Earlier action-value backups considered three n-step choices: n-step Sarsa, the tree-backup algorithm, and n-step Expected Sarsa. They differ in how they handle action choices encountered while looking ahead. Q(σ) places these choices in one framework. The decision can be made separately at each step instead of being forced to remain the same throughout the backup.
| Choice at a look-ahead step | Meaning |
|---|---|
| Sampled action | Use the action that was sampled at that step, as in the Sarsa-style choice. |
| Expected action values | Account for possible actions through an expectation, as in the tree-backup-style choice. |
| Mixed choice | Allow the decision to differ from one step to another within the same backup. |
The action-selection alternatives organized by n-step Q(σ).
Reading σt as Sampling Degree
The symbol σt describes the degree of sampling used at step t. When σt is 1, the step uses full sampling. When σt is 0, the step uses pure expectation. Values between 0 and 1 represent an intermediate degree of sampling.
The subscript t matters because the degree of sampling can vary from one step to another. The source also allows σt to be set as a function of the state, the action, or the state-action pair at time t. Therefore, one n-step backup can contain different sampling choices at different points.
Interpreting a Three-Step Pattern
Consider an n-step backup whose three successive decisions use σ values of 1, 0, and an intermediate value.
First step: σ at the first step is 1, so this step uses full sampling.
Second step: σ at the second step is 0, so this step uses a pure expectation.
Third step: The intermediate σ value represents a degree of sampling between the fully sampled and pure expected cases.
Q(σ) permits the backup to use different sampling degrees at different steps rather than one fixed choice everywhere.
Recovering Earlier Backup Methods
The unifying claim of Q(σ) is that earlier methods can be described through particular patterns of sampling choices. A fully sampled pattern gives the Sarsa-style behavior. A fully expected pattern gives the tree-backup-style behavior. Expected Sarsa is represented by the expected-action choice in the corresponding one-step or n-step backup. Mixed patterns are also allowed, so the framework can move between these established choices instead of treating them as unrelated methods.
Mistakes in Direction and Sampling
Recomputing every backup immediately after one value changes.
Backward focusing selects the actions that lead directly into the changed state as the immediate useful backups.
Fix:
Start with direct incoming actions, then follow predecessor states only when their values may change.Assuming backward focusing only applies to an explicit goal state.
The method is general: a changed value can focus planning even when there is no single explicit goal state.
Fix:
Ask which predecessor states and actions may depend on the changed value.Treating σt as one fixed choice for the whole backup.
The subscript t indicates that the degree of sampling can vary from step to step.
Fix:
Read each σt separately and allow mixed patterns.Interpreting σt = 0 as full sampling.
σt = 0 represents pure expectation, while σt = 1 represents full sampling.
Fix:
Use the endpoints carefully: zero means pure expectation and one means full sampling.
Practice and Review
A state value has just changed. Explain which actions should receive attention first, what can happen after those backups change predecessor values, and how the answer changes if the backup uses σt = 0 rather than σt = 1 at one look-ahead step.
Hints
- Begin with actions that lead directly into the changed state.
- Then consider predecessor states whose values may change after those incoming actions are updated.
- σt = 0 means pure expectation, while σt = 1 means full sampling.
- Backward focusing starts planning at a changed value and moves toward states that may depend on it. The first useful one-step backups are for actions leading directly into the changed state. If predecessor values then change, propagation can continue backward. This reasoning applies to changed rewards and values even without an explicit goal state. For action-value backups, n-step Q(σ) unifies sampled and expected choices: σt = 1 means full sampling, σt = 0 means pure expectation, and intermediate values allow mixed degrees. Different σ patterns describe Sarsa, tree backup, and Expected Sarsa.
Key Takeaways
- Backward focusing directs planning from a changed state value toward relevant predecessor states.
- Direct incoming actions are the first useful one-step backups after a state value changes.
- Value changes can propagate backward through predecessor states and do not require an explicit goal state.
- n-step Q(σ) chooses between sampled actions and expected action values separately at each step.
- σt = 1 means full sampling, σt = 0 means pure expectation, and particular patterns recover Sarsa, tree backup, and Expected Sarsa.