Concepts / Value-Function Approximation

Value-Function Approximation

Reinforcement learning often generates data during interaction with an environment or with a model, so useful approximation methods need online learning.

  • Programming

Learning While Interaction Continues

In many reinforcement-learning settings, the training data is not prepared once and then handed to the learner as a fixed dataset. The agent generates data while it interacts with an environment, or while it interacts with a model of that environment. This changes what a useful value-function approximation method must do: it must learn as new data arrives.

producesupdatesis revisitedEnvironment or modelongoing interactionStatic datasetavailable before trainingNew examplesarrive incrementallyRepeated revisitingrequired by a static-onlyapproachValue approximationupdates while data arrives
How does data flow from ongoing environment interaction into incremental value-function updates, and how does that differ from training once on a fixed dataset?

Online learning means updating from data as it is acquired rather than relying only on a static dataset. For reinforcement learning, this requirement follows from the way training data commonly arrives during interaction.

A Target That Moves

A target function is nonstationary when it changes over time. In a stationary learning problem, the learner is trying to match one fixed mapping from inputs to desired outputs. In reinforcement learning, the values used as learning targets can change as the learning process changes. The learner is therefore facing two kinds of movement at once: new examples may arrive, and the desired value associated with an example may shift.

earlier mappinglater mappingState AinputTarget value 4earlier targetState Asame inputTarget value 7later target
What changes over time in the mapping from states to target values, and why can the same state receive different learning targets at different moments?

One State, Two Learning Moments

Suppose a value approximation method receives an example for State A early in learning and a later example for the same State A. The target in the first example is 4, while the later target is 7.

First observation: The method receives State A with a target value of 4 and adjusts its approximation toward that target.

Learning process changes: The reinforcement-learning process changes, so the value used as the target for State A is no longer the same.

Later observation: The method receives State A again, now with a target value of 7, and must adapt rather than treating the earlier target as permanently correct.

The input state stayed the same, but its learning target changed. This is an illustrative example of a nonstationary target function.

Why Value Targets Change

Two mechanisms identified in reinforcement learning can make target values change. First, generalized policy iteration can change the policy. When the policy changes, the values associated with states can change as well. Second, dynamic programming and temporal-difference learning can use bootstrapping. In bootstrapping, target values can depend on value estimates that are themselves being updated. As those successor-state estimates change, the target used for another update can change too.

policy change altersbootstrapping altersprovides update targetPolicychangesTarget valuechangesCurrent state valueupdated toward targetSuccessor-stateestimatechanges throughbootstrapping
How does a change in the policy or in a successor state's estimated value alter the target used to update the current value estimate?

Policy Change and Bootstrapping

Consider a current state whose value is being approximated. Its target is affected first by a policy change and later by a change in the estimated value of a successor state.

Initial target: The current state receives a target based on the policy and value estimates available at that moment.

Policy changes: A changed policy can produce a different target value for the current state because the future behavior being evaluated has changed.

Successor estimate changes: If the target uses a successor state's estimated value, bootstrapping means that a change in that estimate can also change the target.

Incremental adaptation: The approximation method must continue updating as these new targets appear instead of assuming that the initial target remains fixed.

Policy changes and bootstrapping provide two distinct reasons that a value target can move during learning.

Assessing an Approximation Method

A value-function approximation method should not be judged only by how accurately it fits a collection of examples. It must also fit the way those examples appear in reinforcement learning. Two evaluation questions are central: Can the method learn incrementally as examples arrive? Can it adapt when the target values shift?

evaluateevaluatecheck for dependenceApproximationmethodcandidate for valuelearningIncremental learninglearns as data arrivesTarget adaptationresponds to changing valuesStatic-datasetdependencerelies only on fixed data
How do approximation methods differ when examples arrive incrementally and target values shift?
QuestionWhat it testsWhy it matters
Can the method learn incrementally?Whether it updates from examples as they are acquiredReinforcement-learning data commonly arrives during interaction with an environment or a model
Can the method adapt to changing targets?Whether it can respond when the target function or target values changePolicy changes and bootstrapping can make the desired values move
Does it only fit a static collection of examples?Whether it depends on repeatedly revisiting a fixed datasetStatic-dataset fitting alone does not address incremental data arrival

Evaluation criteria for value-function approximation methods

Many familiar supervised-learning methods can, in principle, approximate value functions from backup-generated examples. The source lists artificial neural networks, decision trees, and several forms of multivariate regression. However, describing a problem as supervised learning does not by itself show that every such method is equally suitable for reinforcement learning. Suitability depends on whether the method matches both the incremental data stream and the changing targets.

Common Evaluation Mistakes

  • Treating online learning and nonstationarity as the same idea.

    Online learning concerns the arrival of data, while nonstationarity concerns changes in the target function or target values.

    Fix: Evaluate both requirements separately.

  • Assuming that accurate fitting on a fixed dataset is sufficient.

    Reinforcement-learning data commonly arrives incrementally, and the targets can change during learning.

    Fix: Ask whether the method can update during ongoing interaction and adapt to shifting targets.

  • Assuming that a supervised-learning formulation guarantees suitability.

    These methods can, in principle, be used to approximate value functions, but the learning problem's data-arrival pattern and changing targets still matter.

    Fix: Assess the method against incremental learning and target adaptation.

  • Overlooking bootstrapping as a source of target change.

    Bootstrapping can make the target change when the successor-state estimate changes.

    Fix: Track whether the values used to construct targets are themselves being updated.

Method-Selection Practice

MEDIUM

You are comparing two candidate value-function approximation methods. Method A fits a fixed collection of examples accurately but is intended to rely on repeatedly revisiting that collection. Method B updates as new examples arrive and can change its approximation when target values shift. Which method better matches the requirements described for reinforcement learning, and what two properties support your choice?

Hints
  • Separate the question of how data arrives from the question of whether targets change.
  • Look for incremental learning and adaptation to nonstationary targets.

What do you think happens?

Before reading the explanation, choose the stronger match for ongoing reinforcement-learning data.

  • Method A, because fixed-dataset accuracy is the only requirement
  • Method B, because it learns incrementally and responds to changing targets
Reveal answer

Answer: Method B, because it learns incrementally and responds to changing targets

Reinforcement learning commonly produces examples during interaction, so online learning is needed. Policy changes and bootstrapping can also make target values change, so adaptation to nonstationary targets is another requirement.

Key Takeaways

  1. Reinforcement learning commonly generates training data during interaction with an environment or a model, so useful approximation methods need online learning.
  2. Online learning means updating from data as it is acquired rather than relying only on a static dataset.
  3. A target function is nonstationary when it changes over time.
  4. Policy changes and bootstrapping can cause the values used as learning targets to change.
  5. A value-function approximation method should be evaluated for both incremental learning and adaptation to changing targets, not only for accuracy on a fixed collection of examples.

Key Takeaways

  • Value-function approximation in reinforcement learning must match an incremental stream of data.
  • Online learning addresses how examples arrive, while nonstationarity addresses how target values change.
  • Policy changes and bootstrapping can both alter the targets used to update value estimates.
  • A suitable approximation method must learn incrementally and adapt as its targets shift.