Concepts / Control with Function Approximation

Control with Function Approximation

Function approximation introduces distinctive challenges into reinforcement learning.

  • Programming

Why Approximation Changes Control

Function approximation introduces distinctive challenges into reinforcement learning. The central difficulty is not merely estimating something with less precision. In this setting, the way learning is organized can create issues that do not normally arise in conventional supervised learning.

A useful way to study the topic is to follow the path from a state representation to an estimated value or policy, and then from that estimate to a control decision. The important question is not only whether the estimate is accurate. We must also ask whether the learning target changes, whether later estimates depend on earlier estimates, and whether the objective being optimized is still the right one.

From States to Decisions

In a conceptual function-approximation setup, many possible states are represented through an estimated value or policy rather than being treated as isolated learning entries. The estimate then contributes to a control decision. This shared representation is useful for understanding why information associated with one state can affect estimates or decisions associated with other states.

inputproducesinformsState representationcurrent situationFunction approximatorshared estimateEstimated value orpolicyapproximate resultControl actionselected response
How does a function approximator transform a state representation into an estimated value or action?

A shared estimate in a conceptual control task

Suppose a controller encounters two states that have similar representations. Learning from one state changes the shared estimate used for both states. What consequence should you examine?

Represent the states: The two situations are supplied to the function approximator through their state representations.

Produce estimates: The approximator produces estimated values or policy information for the situations rather than requiring a separate exact entry for every possible situation.

Update one case: An update associated with one state can influence the shared estimate and therefore affect the estimate for another similar state.

Reconsider the action: Because the estimated values or policy information may have changed, the controller may reconsider which action to select.

The key issue is generalization: learning connected with one state can affect other similar states. In control, that influence matters because changed estimates can change later decisions.

learning updateshared influenceState Aestimate for AState Aupdated estimateState Bestimate for BState Baffected estimate
How can learning from one state change the estimated values or actions for other similar states?

Three Separate Learning Difficulties

The source identifies nonstationarity, bootstrapping, and delayed targets as separate issues. This separation is important because each describes a different reason that learning can become difficult when function approximation is used in reinforcement learning.

IssueWhat to keep separateWhy it matters for analysis
NonstationarityThe learning situation may not remain fixedAn estimate can be trained while the effective problem it faces is changing
BootstrappingLearning can rely on other estimatesAn estimate may be influenced by estimates rather than only by a final observed outcome
Delayed targetsThe learning signal may arrive laterThe outcome used to judge an earlier decision may not be immediately available

The three challenges should be analyzed as distinct issues rather than collapsed into one label.

These issues can combine. For example, a changing learning situation can affect the estimates used as learning information, while delayed outcomes make it harder to determine which earlier decision should receive credit. The important discipline is to identify which issue is being discussed at each step instead of assuming that every difficulty has the same cause.

When Estimates Affect Later Control

Control creates a feedback relationship between estimates and later behavior. An estimate informs a decision; that decision contributes to what is observed next; and later learning can change the estimate again. With function approximation, an update can also influence estimates for related states. This makes it necessary to examine not only the current estimate, but also how the estimate changes future decisions.

informscontributes toupdateschangescontinues feedbackFunction estimatecurrent estimateControl decisionaction selectionLater outcomedelayed informationUpdated estimatechanged estimateFuture decisionnew action selection
How can an updated function estimate alter later decisions and create oscillation or divergence during control?
representsinfluences estimatesState-action entriesindividual representationFunction parametersshared representationEntry updatelocalized viewEstimate updatepossible shared influence
What changes when individual state-action entries are replaced by shared function parameters?

Approximation Errors in Action Selection

A control decision depends on the estimate available to the controller. If that estimate is inaccurate, the controller can select an action that differs from the action it would select under a more accurate estimate. This is why approximation error is especially important in control: an error is not only a measurement problem; it can alter subsequent behavior.

estimatedinfluencesmay differ fromStaterepresentationcurrent situationInaccurate estimateapproximation errorSelected actionbased on estimateOptimal actioncomparison point
How can an inaccurate estimated value cause the controller to select a different action from the optimal one?

Checking the decision, not only the estimate

A controller has two available actions in a state. Its approximate estimates rank Action A above Action B, but a more accurate assessment would rank Action B above Action A. What should the learner identify?

Locate the approximation: The controller is relying on an estimate rather than an exact, guaranteed assessment.

Compare the rankings: The approximate estimate ranks the actions differently from the more accurate assessment.

Identify the control consequence: The estimate changes the action selected by the controller.

Connect the consequence to learning: Because the selected action affects later experience, the initial approximation error can become part of the continuing control-and-learning process.

In approximate control, evaluate both estimation quality and the decisions produced by the estimate.

Practice the Diagnostic View

MEDIUM

A learning system changes its estimated value after receiving information that arrives later. The updated estimate changes a future control decision, and the changed decision affects the information received afterward. Identify which of the three named challenges are potentially involved, and explain why you should analyze them separately.

Hints
  • Look first for the fact that information arrives later.
  • Then ask whether the update relies on an existing estimate.
  • Finally ask whether the learning situation or effective target changes as control decisions change.
  • Delayed targets may be involved because the information used for learning arrives after the earlier decision.
  • Bootstrapping may be involved when an estimate is influenced by another estimate.
  • Nonstationarity may be involved when the learning situation does not remain fixed as behavior and estimates change.
  • The three labels should remain distinct even when one example contains all three.
  • Treating function approximation as only a precision issue

    The source emphasizes distinctive issues in reinforcement learning that do not normally arise in conventional supervised learning.

    Fix: Examine how approximation interacts with the learning process, targets, and control decisions.

  • Combining nonstationarity, bootstrapping, and delayed targets into one challenge

    The source explicitly identifies these as separate issues.

    Fix: Name the specific feature present in the situation and then describe any interactions separately.

  • Assuming that the original learning objectives remain unchanged

    The source states that the learning objectives must be reconsidered in this setting.

    Fix: Re-examine the objective instead of assuming it transfers unchanged.

Key Takeaways

  1. Function approximation introduces distinctive challenges into reinforcement learning.
  2. Nonstationarity, bootstrapping, and delayed targets are separate issues that should be analyzed separately.
  3. Approximation can affect more than an estimate: it can change control decisions and later learning.
  4. Learning objectives must be reconsidered rather than assumed to remain unchanged.
  5. The broader progression includes on-policy prediction and control, off-policy methods, eligibility traces, and policy-gradient methods.

Key Takeaways

  • Function approximation creates reinforcement-learning challenges that do not normally arise in conventional supervised learning.
  • Nonstationarity, bootstrapping, and delayed targets are distinct issues, even when they appear together.
  • Shared estimates can influence related states and change later control decisions.
  • Approximation errors matter because they can alter behavior, not only numerical estimates.
  • The learning objectives should be re-examined in the function-approximation setting.