Value-Function Approximation
Reinforcement learning often generates data during interaction with an environment or with a model, so useful approximation methods need online learning.
Learning While Interaction Continues
In many reinforcement-learning settings, the training data is not prepared once and then handed to the learner as a fixed dataset. The agent generates data while it interacts with an environment, or while it interacts with a model of that environment. This changes what a useful value-function approximation method must do: it must learn as new data arrives.
Online learning means updating from data as it is acquired rather than relying only on a static dataset. For reinforcement learning, this requirement follows from the way training data commonly arrives during interaction.
A Target That Moves
A target function is nonstationary when it changes over time. In a stationary learning problem, the learner is trying to match one fixed mapping from inputs to desired outputs. In reinforcement learning, the values used as learning targets can change as the learning process changes. The learner is therefore facing two kinds of movement at once: new examples may arrive, and the desired value associated with an example may shift.
One State, Two Learning Moments
Suppose a value approximation method receives an example for State A early in learning and a later example for the same State A. The target in the first example is 4, while the later target is 7.
First observation: The method receives State A with a target value of 4 and adjusts its approximation toward that target.
Learning process changes: The reinforcement-learning process changes, so the value used as the target for State A is no longer the same.
Later observation: The method receives State A again, now with a target value of 7, and must adapt rather than treating the earlier target as permanently correct.
The input state stayed the same, but its learning target changed. This is an illustrative example of a nonstationary target function.
Why Value Targets Change
Two mechanisms identified in reinforcement learning can make target values change. First, generalized policy iteration can change the policy. When the policy changes, the values associated with states can change as well. Second, dynamic programming and temporal-difference learning can use bootstrapping. In bootstrapping, target values can depend on value estimates that are themselves being updated. As those successor-state estimates change, the target used for another update can change too.
Policy Change and Bootstrapping
Consider a current state whose value is being approximated. Its target is affected first by a policy change and later by a change in the estimated value of a successor state.
Initial target: The current state receives a target based on the policy and value estimates available at that moment.
Policy changes: A changed policy can produce a different target value for the current state because the future behavior being evaluated has changed.
Successor estimate changes: If the target uses a successor state's estimated value, bootstrapping means that a change in that estimate can also change the target.
Incremental adaptation: The approximation method must continue updating as these new targets appear instead of assuming that the initial target remains fixed.
Policy changes and bootstrapping provide two distinct reasons that a value target can move during learning.
Assessing an Approximation Method
A value-function approximation method should not be judged only by how accurately it fits a collection of examples. It must also fit the way those examples appear in reinforcement learning. Two evaluation questions are central: Can the method learn incrementally as examples arrive? Can it adapt when the target values shift?
| Question | What it tests | Why it matters |
|---|---|---|
| Can the method learn incrementally? | Whether it updates from examples as they are acquired | Reinforcement-learning data commonly arrives during interaction with an environment or a model |
| Can the method adapt to changing targets? | Whether it can respond when the target function or target values change | Policy changes and bootstrapping can make the desired values move |
| Does it only fit a static collection of examples? | Whether it depends on repeatedly revisiting a fixed dataset | Static-dataset fitting alone does not address incremental data arrival |
Evaluation criteria for value-function approximation methods
Many familiar supervised-learning methods can, in principle, approximate value functions from backup-generated examples. The source lists artificial neural networks, decision trees, and several forms of multivariate regression. However, describing a problem as supervised learning does not by itself show that every such method is equally suitable for reinforcement learning. Suitability depends on whether the method matches both the incremental data stream and the changing targets.
Common Evaluation Mistakes
Treating online learning and nonstationarity as the same idea.
Online learning concerns the arrival of data, while nonstationarity concerns changes in the target function or target values.
Fix:
Evaluate both requirements separately.Assuming that accurate fitting on a fixed dataset is sufficient.
Reinforcement-learning data commonly arrives incrementally, and the targets can change during learning.
Fix:
Ask whether the method can update during ongoing interaction and adapt to shifting targets.Assuming that a supervised-learning formulation guarantees suitability.
These methods can, in principle, be used to approximate value functions, but the learning problem's data-arrival pattern and changing targets still matter.
Fix:
Assess the method against incremental learning and target adaptation.Overlooking bootstrapping as a source of target change.
Bootstrapping can make the target change when the successor-state estimate changes.
Fix:
Track whether the values used to construct targets are themselves being updated.
Method-Selection Practice
You are comparing two candidate value-function approximation methods. Method A fits a fixed collection of examples accurately but is intended to rely on repeatedly revisiting that collection. Method B updates as new examples arrive and can change its approximation when target values shift. Which method better matches the requirements described for reinforcement learning, and what two properties support your choice?
Hints
- Separate the question of how data arrives from the question of whether targets change.
- Look for incremental learning and adaptation to nonstationary targets.
What do you think happens?
Before reading the explanation, choose the stronger match for ongoing reinforcement-learning data.
Reveal answer
Answer: Method B, because it learns incrementally and responds to changing targets
Reinforcement learning commonly produces examples during interaction, so online learning is needed. Policy changes and bootstrapping can also make target values change, so adaptation to nonstationary targets is another requirement.
Key Takeaways
- Reinforcement learning commonly generates training data during interaction with an environment or a model, so useful approximation methods need online learning.
- Online learning means updating from data as it is acquired rather than relying only on a static dataset.
- A target function is nonstationary when it changes over time.
- Policy changes and bootstrapping can cause the values used as learning targets to change.
- A value-function approximation method should be evaluated for both incremental learning and adaptation to changing targets, not only for accuracy on a fixed collection of examples.
Key Takeaways
- Value-function approximation in reinforcement learning must match an incremental stream of data.
- Online learning addresses how examples arrive, while nonstationarity addresses how target values change.
- Policy changes and bootstrapping can both alter the targets used to update value estimates.
- A suitable approximation method must learn incrementally and adapt as its targets shift.