Control with Function Approximation
Function approximation introduces distinctive challenges into reinforcement learning.
Why Approximation Changes Control
Function approximation introduces distinctive challenges into reinforcement learning. The central difficulty is not merely estimating something with less precision. In this setting, the way learning is organized can create issues that do not normally arise in conventional supervised learning.
A useful way to study the topic is to follow the path from a state representation to an estimated value or policy, and then from that estimate to a control decision. The important question is not only whether the estimate is accurate. We must also ask whether the learning target changes, whether later estimates depend on earlier estimates, and whether the objective being optimized is still the right one.
From States to Decisions
In a conceptual function-approximation setup, many possible states are represented through an estimated value or policy rather than being treated as isolated learning entries. The estimate then contributes to a control decision. This shared representation is useful for understanding why information associated with one state can affect estimates or decisions associated with other states.
A shared estimate in a conceptual control task
Suppose a controller encounters two states that have similar representations. Learning from one state changes the shared estimate used for both states. What consequence should you examine?
Represent the states: The two situations are supplied to the function approximator through their state representations.
Produce estimates: The approximator produces estimated values or policy information for the situations rather than requiring a separate exact entry for every possible situation.
Update one case: An update associated with one state can influence the shared estimate and therefore affect the estimate for another similar state.
Reconsider the action: Because the estimated values or policy information may have changed, the controller may reconsider which action to select.
The key issue is generalization: learning connected with one state can affect other similar states. In control, that influence matters because changed estimates can change later decisions.
Three Separate Learning Difficulties
The source identifies nonstationarity, bootstrapping, and delayed targets as separate issues. This separation is important because each describes a different reason that learning can become difficult when function approximation is used in reinforcement learning.
| Issue | What to keep separate | Why it matters for analysis |
|---|---|---|
| Nonstationarity | The learning situation may not remain fixed | An estimate can be trained while the effective problem it faces is changing |
| Bootstrapping | Learning can rely on other estimates | An estimate may be influenced by estimates rather than only by a final observed outcome |
| Delayed targets | The learning signal may arrive later | The outcome used to judge an earlier decision may not be immediately available |
The three challenges should be analyzed as distinct issues rather than collapsed into one label.
These issues can combine. For example, a changing learning situation can affect the estimates used as learning information, while delayed outcomes make it harder to determine which earlier decision should receive credit. The important discipline is to identify which issue is being discussed at each step instead of assuming that every difficulty has the same cause.
When Estimates Affect Later Control
Control creates a feedback relationship between estimates and later behavior. An estimate informs a decision; that decision contributes to what is observed next; and later learning can change the estimate again. With function approximation, an update can also influence estimates for related states. This makes it necessary to examine not only the current estimate, but also how the estimate changes future decisions.
Approximation Errors in Action Selection
A control decision depends on the estimate available to the controller. If that estimate is inaccurate, the controller can select an action that differs from the action it would select under a more accurate estimate. This is why approximation error is especially important in control: an error is not only a measurement problem; it can alter subsequent behavior.
Checking the decision, not only the estimate
A controller has two available actions in a state. Its approximate estimates rank Action A above Action B, but a more accurate assessment would rank Action B above Action A. What should the learner identify?
Locate the approximation: The controller is relying on an estimate rather than an exact, guaranteed assessment.
Compare the rankings: The approximate estimate ranks the actions differently from the more accurate assessment.
Identify the control consequence: The estimate changes the action selected by the controller.
Connect the consequence to learning: Because the selected action affects later experience, the initial approximation error can become part of the continuing control-and-learning process.
In approximate control, evaluate both estimation quality and the decisions produced by the estimate.
Practice the Diagnostic View
A learning system changes its estimated value after receiving information that arrives later. The updated estimate changes a future control decision, and the changed decision affects the information received afterward. Identify which of the three named challenges are potentially involved, and explain why you should analyze them separately.
Hints
- Look first for the fact that information arrives later.
- Then ask whether the update relies on an existing estimate.
- Finally ask whether the learning situation or effective target changes as control decisions change.
- Delayed targets may be involved because the information used for learning arrives after the earlier decision.
- Bootstrapping may be involved when an estimate is influenced by another estimate.
- Nonstationarity may be involved when the learning situation does not remain fixed as behavior and estimates change.
- The three labels should remain distinct even when one example contains all three.
Treating function approximation as only a precision issue
The source emphasizes distinctive issues in reinforcement learning that do not normally arise in conventional supervised learning.
Fix:
Examine how approximation interacts with the learning process, targets, and control decisions.Combining nonstationarity, bootstrapping, and delayed targets into one challenge
The source explicitly identifies these as separate issues.
Fix:
Name the specific feature present in the situation and then describe any interactions separately.Assuming that the original learning objectives remain unchanged
The source states that the learning objectives must be reconsidered in this setting.
Fix:
Re-examine the objective instead of assuming it transfers unchanged.
Key Takeaways
- Function approximation introduces distinctive challenges into reinforcement learning.
- Nonstationarity, bootstrapping, and delayed targets are separate issues that should be analyzed separately.
- Approximation can affect more than an estimate: it can change control decisions and later learning.
- Learning objectives must be reconsidered rather than assumed to remain unchanged.
- The broader progression includes on-policy prediction and control, off-policy methods, eligibility traces, and policy-gradient methods.
Key Takeaways
- Function approximation creates reinforcement-learning challenges that do not normally arise in conventional supervised learning.
- Nonstationarity, bootstrapping, and delayed targets are distinct issues, even when they appear together.
- Shared estimates can influence related states and change later control decisions.
- Approximation errors matter because they can alter behavior, not only numerical estimates.
- The learning objectives should be re-examined in the function-approximation setting.