Understanding Backups in Value Prediction
A backup is also a specification of desired value-function behavior.
Why a Backup Matters
A backup is not merely an instruction to adjust a stored number. It also specifies desired behavior for a value function. It identifies a state and indicates the numerical output that the value function should move toward for that state.
The central interpretation is: a backup tells the value function what output is desired when the input is a particular state.
Reading the Backup Pair
The notation s ↦→ g represents a state paired with a desired numerical output. Read it as an input-output example: the input is state s, and the desired output is the number g. The estimated value for s should move toward g.
Interpreting One Backup
Suppose a backup is written as state_A ↦→ 10.
Identify the input: The input to the value function is state_A.
Identify the desired output: The desired numerical output paired with state_A is 10.
Interpret the update: The estimated value for state_A should move toward 10. The backup therefore describes desired behavior for the function at that input.
state_A ↦→ 10 is an input-output training example: state_A is the input and 10 is the desired output.
Desired Function Behavior
The backup can be viewed as a specification of what the value function should do for a particular state. Before using the backup, the function has an existing estimated output for that state. The backup supplies a desired number, and the value estimate should move toward that number.
From Backups to Function Approximation
Function approximation uses backup pairs as training examples for value prediction. Function approximation methods expect examples that describe the desired behavior of the function being approximated. In this setting, the backup supplies the example, and the resulting approximate function is interpreted as the estimated value function.
| Implementation viewpoint | What the backup does |
|---|---|
| Table of state values | Shift the table entry for s partway toward g and leave the other state entries unchanged. |
| Function approximation | Use the backup as evidence about the input-output behavior the approximate function should learn. |
This change in viewpoint is important. A table can directly adjust one stored entry. An approximation method is not restricted to that representation, so one update can change estimated values for many states as part of the same update. It need not change only the estimate for state_B because the approximator represents value behavior across states rather than treating every state as an isolated table entry.
Learning Versus Prediction
Supervised learning and the estimated value function have different roles. Supervised learning is the machine-learning approach of learning from input-output examples. In value prediction, backups provide those examples. The estimated value function is the approximate function produced after the method processes such examples; it is then interpreted as the value function that makes the predictions.
Supervised learning describes how the function is learned. The estimated value function is the resulting approximate function used to produce value estimates.
Common Interpretation Errors
Treating s ↦→ g as only a command to replace one stored number.
A backup also specifies desired input-output behavior for the value function.
Fix:
Interpret s as the input state and g as the desired numerical output toward which the estimate should move.Assuming that function approximation must change only the estimate for the backed-up state.
Function approximation can change estimated values for many states as part of the same update.
Fix:
Remember that the backup meaning stays the same while the approximation method can implement it through a broader function update.Confusing supervised learning with the estimated value function.
Supervised learning is the approach for learning from examples, whereas the estimated value function is the resulting approximate function.
Fix:
Identify backups as training examples, supervised learning as the learning approach, and the estimated value function as the learned approximate function.
When reading a backup, ask two questions: what is the input state, and what numerical output should the value function move toward? These questions preserve both the update meaning and the function-behavior meaning.
Check Your Interpretation
A backup is written as state_C ↦→ 4. Explain what state_C and 4 represent, how this pair can be used by function approximation, and why an approximation update need not change only the estimate for state_C.
Hints
- Identify the input and the desired numerical output.
- Connect the pair to input-output training examples.
- Contrast a direct table entry update with a function-approximation update.
What do you think happens?
If a table of state values implements a backup for state_C, must the entries for every other state change?
Reveal answer
Answer: No, the other state entries are left unchanged.
For a table of state values, implementing a backup can shift the table entry for the backed-up state partway toward the desired output while leaving the other state entries unchanged.
Key Takeaways
- A backup specifies desired value-function behavior for a state.
- The notation s ↦→ g pairs an input state with a desired numerical output.
- Function approximation treats backup pairs as input-output training examples for value prediction.
- A table can adjust one state entry while leaving other entries unchanged; an approximation update can change estimated values for many states.
- Supervised learning is the approach that learns from examples, while the estimated value function is the resulting approximate function that produces value estimates.
Key Takeaways
- A backup tells a value function what output is desired for a particular state.
- The pair s ↦→ g is an input-output training example.
- Function approximation uses backup examples to learn an approximate value function.
- A function-approximation update can affect estimated values for many states, unlike a direct table-entry adjustment.
- Supervised learning is the learning approach; the estimated value function is the resulting prediction-making function.