Nonstationary Target Functions
Backups provide training examples for function approximation in value prediction.
Why a Backup Looks Like a Training Example
A reinforcement learning backup can be viewed as a conventional training example for value prediction. The learning system has an experience from the environment or a model, uses a backup procedure to determine a target, and adjusts a function approximator so that its value prediction can learn from that example.
This observation connects reinforcement learning with supervised learning. Artificial neural networks, decision trees, and several forms of multivariate regression are all possible function approximation methods. However, the connection is only a starting point. A method can be valid for supervised learning and still be a poor fit for reinforcement learning because reinforcement learning usually does not provide one fixed dataset of unchanged examples.
The Training Situation in Reinforcement Learning
In a fixed supervised-learning setting, a method may be designed to receive one dataset and make repeated passes over examples that do not change. Reinforcement learning commonly works differently. An agent gathers examples gradually while interacting with an environment or with a model of the environment.
An Incrementally Built Training Stream
Compare a fixed supervised-learning dataset with the way reinforcement learning commonly gathers backup examples.
Fixed-data setting: A supervised-learning method may receive a fixed set of examples and repeatedly process those unchanged examples.
First interaction: In reinforcement learning, an agent interacts with an environment or model and obtains information from which a backup example can be formed.
Later interaction: Additional interaction produces additional examples. The training data therefore arrives incrementally rather than being available as one unchanged dataset from the beginning.
Method-selection consequence: A method that is valid for supervised learning may still be a poor reinforcement-learning choice if it does not learn efficiently while new data is arriving.
Online learning matters because reinforcement learning commonly receives examples gradually during interaction.
How Bootstrapping Moves the Target
A target is nonstationary when the target function changes over time. One source of this change is bootstrapping. In bootstrapping methods such as dynamic programming and temporal-difference learning, values used in later backups can depend on current value estimates. When those estimates change, the values used to form subsequent targets can change as well.
A Backup Target That Changes with the Estimate
Suppose a learner uses its current value estimates when forming a backup target. What happens to a later training example after the learner's estimates have changed?
Initial estimate: The learner forms a backup using the value estimates currently available.
Learning update: The function approximator learns from the backup example, so its value predictions change.
Later backup: A subsequent backup may use the changed value predictions. The resulting target is therefore not necessarily the same as the earlier target.
Interpretation: The learner is not merely repeating training on a permanently fixed target function. Its own changing estimates can help determine later targets.
Bootstrapping can make the target function change as learning proceeds.
How Policy Changes Alter Values
Changing policies are another source of nonstationary target functions. A value prediction depends on the future outcomes associated with the situation being evaluated. When the policy changes, the future outcomes considered by the learning process can change, so the target value assigned to the same state can change as well.
This means that changing targets do not arise only because a function approximator is being updated. The behavior policy itself can change the outcomes represented by value predictions. A method suited to reinforcement learning must therefore be considered in a setting where both new data and changing target behavior may be present.
Choosing an Approximator for the Situation
Function approximation produces value predictions from backup examples. The approximation method is not identical to the value function: the value function is what the learning process seeks to predict, while the approximation method is the mechanism used to learn those predictions.
| Question | Fixed supervised-learning setting | Reinforcement-learning setting |
|---|---|---|
| How does data arrive? | A fixed dataset may be available for repeated passes. | Examples commonly arrive incrementally during interaction with an environment or model. |
| Do targets stay the same? | The examples may remain unchanged during training. | Targets can change because of bootstrapping and changing policies. |
| What should method selection ask? | Can the method learn from supervised examples? | Can it learn while new data arrives and while the target function changes? |
| What does availability mean? | The method is a possible supervised-learning method. | The method may or may not be well suited to reinforcement learning. |
Assuming that every supervised-learning method is equally suitable for value prediction in reinforcement learning.
Reinforcement learning commonly provides examples incrementally, and its targets can change over time.
Fix:
Check both online-learning ability and the ability to continue learning when the target function changes.Treating a backup target as permanently fixed.
Bootstrapping can use changing value estimates, and changing policies can alter the future outcomes represented by targets.
Fix:
Analyze how estimates and policies may change the targets used by later backups.Confusing the value function with the function approximation method.
The value function is what the learning process seeks to predict; the approximation method is the mechanism used to learn predictions from backup examples.
Fix:
Keep the predicted object and the learning mechanism conceptually separate.
Check Your Understanding
A function approximation method can learn from supervised examples, but it expects one fixed dataset and repeated passes over unchanged examples. Explain why that fact alone does not show that the method is well suited to reinforcement learning.
Hints
- Identify how reinforcement-learning examples commonly arrive.
- Identify two reasons that reinforcement-learning targets can change.
- Separate being a candidate supervised-learning method from being suitable for the reinforcement-learning training situation.
What do you think happens?
A learner changes its value estimates after one backup. Should you expect every later bootstrapped target to remain identical to the earlier target?
Reveal answer
Answer: No. Later targets can change because bootstrapping methods can use the learner's changing value estimates.
Bootstrapping is one source of nonstationarity. Changing policies provide another source because they can change the future outcomes represented by the target value.
Key Takeaways
- Each reinforcement learning backup can be treated as a training example for value prediction.
- Reinforcement learning commonly gathers examples incrementally during interaction with an environment or model.
- Bootstrapping can make targets change because later backups may use updated value estimates.
- Changing policies can alter future outcomes and therefore alter target values for the same state.
- A method can be usable for supervised learning without being well suited to reinforcement learning; method selection must consider incremental data and changing targets.
Key Takeaways
- Backups turn reinforcement-learning experience into training examples for a value-prediction method.
- Online learning is important because examples commonly arrive gradually during interaction.
- Nonstationary targets arise from bootstrapping and changing policies.
- Supervised-learning compatibility is not enough; a suitable method must also handle incremental data and changing target functions.