Concepts / Approximate Solutions for Large State Spaces

Approximate Solutions for Large State Spaces

Function approximation turns examples of a desired function into a broader estimate.

  • Programming

From Limited Experience to a Broader Estimate

A reinforcement-learning system may have only some examples of the function it needs. The central problem is not simply how to store those examples, but how to use them to estimate the function more broadly. Function approximation addresses this problem by turning examples of a desired function into a broader estimate.

The defining operation is generalization: using available examples to represent more of a desired function than the examples themselves directly describe.

paired withsupportsestimatesExample statesselected state examplesFunction estimatebroader representationDesired valuesvalues associated withexamplesUnseen statesestimated by generalization
How does an approximation use values from a limited set of example states to estimate values for many unseen states?

The Boundary Between Examples and Approximation

The most important distinction is between input evidence and the result built from that evidence. Examples associated with a desired function are the available evidence. The approximator uses those examples to construct an estimate intended to represent the entire function. Therefore, function approximation is not merely example storage. It is an attempt to generalize beyond the isolated examples.

Estimating a Value Function

Suppose a reinforcement-learning system has examples for several states, but not for every state it may encounter. How does function approximation change the role of those examples?

Collect examples: The system begins with some examples associated with a desired function. In reinforcement learning, that desired function can be a value function.

Use the examples as evidence: The examples are not treated as isolated facts that answer only the states already observed. They provide information from which a broader estimate can be constructed.

Construct an approximation: An approximator generalizes from the available examples and produces an estimate intended to represent the desired function as a whole.

Estimate beyond the examples: The resulting approximation can represent more of the desired function than the limited collection of examples directly contains.

Function approximation extends limited examples into a broader estimate of a desired function, such as a value function.

retainsconstructsExample storageisolated examplesKnown exampleslimited coverageFunctionapproximationgeneralized estimateBroader estimaterepresentation of thefunction
What changes when reinforcement learning stores a compact approximating function instead of one value for every possible state?

Function Approximation as Supervised Learning

Function approximation is an instance of supervised learning used within reinforcement-learning algorithms. The available examples provide evidence about the desired function, and the approximation process uses that evidence to produce an estimated function.

providesprovidesproducesState examplesinput examplesSupervised learninguses examples as evidenceEstimated functionbroader approximationDesired valuesfunction-related outputs
How do state examples and desired output values flow through a supervised learning process to produce an estimated function?

This perspective clarifies the roles in the process. States and their associated desired values form the examples. Supervised learning describes the use of those examples to construct an estimate. The estimated function is the broader result that can be used within a reinforcement-learning algorithm.

How Generalization Spreads Information

Generalization means that learning from available examples contributes to a broader estimate rather than affecting only the exact examples that were observed. In this sense, examples can support estimates for other states that are related to the available evidence. The key result is broader representation of the desired function.

hassupportsextends toObserved stateavailable exampleFunction evidenceassociated desired valueGeneralizationbuilds broader estimateRelated statesestimated function values
How can learning from one state change the estimated values of other related states?

An approximation should not be described as if every possible state had been directly observed. The source emphasizes that the system may have only some examples. The approximation is valuable precisely because it attempts to represent more of the desired function than those examples explicitly cover.

Choosing an Approximation Method

Many fields provide methods that could serve as function approximators in reinforcement learning. The source identifies machine learning, artificial neural networks, pattern recognition, and statistical curve fitting as possible sources of these methods.

can providecan providecan providecan provideRL functionapproximationdesired function estimateMachine learningpossible methodsArtificial neuralnetworkspossible methodsPattern recognitionpossible methodsStatistical curvefittingpossible methods
How do methods from several fields fit into the common task of approximating a reinforcement-learning function?

The theoretical range is broad: methods from any of these fields can take the role of an approximator within a reinforcement-learning algorithm. However, theoretical possibility does not guarantee practical suitability. A useful choice separates two questions: whether a method can serve as a function approximator in principle, and how naturally it fits inside the particular reinforcement-learning algorithm.

QuestionWhat it evaluates
Can the method approximate a function?Whether it can serve as a function approximator in principle.
Does it fit the reinforcement-learning algorithm?Whether the method naturally fits the particular algorithm in practice.

Separating theoretical possibility from practical fit

Common Misunderstandings

  • Treating function approximation as simple storage of observed examples.

    The defining operation is generalization. The goal is to use examples to represent more of the desired function.

    Fix: Describe the examples as evidence used to construct a broader estimate.

  • Assuming that the desired function must be completely known before approximation begins.

    The source specifically describes the case in which a system has only some examples of the function it needs.

    Fix: Focus on how limited examples can support an estimate of the function as a whole.

  • Confusing a possible method with a universally suitable method.

    The source distinguishes whether a method can serve as an approximator in principle from how naturally it fits a particular algorithm.

    Fix: Evaluate both the method's ability to approximate and its fit with the specific algorithm.

  • Ignoring the supervised-learning aspect.

    The source identifies function approximation as an instance of supervised learning used within reinforcement-learning algorithms.

    Fix: Recognize the examples and associated desired values as the evidence used in the supervised-learning process.

Practice: Trace the Information

MEDIUM

A reinforcement-learning system has examples associated with a value function for only some states. Explain the role of those examples, the operation that turns them into a broader estimate, and why the result should not be called simple example storage.

Hints
  • Identify whether the examples are the final result or the input evidence.
  • Use the term that describes extending information beyond isolated examples.
  • Connect the process to supervised learning.

Practice Answer

Explain how limited value-function examples can support estimates for more states.

Identify the evidence: The examples are associated with the desired function, which can be a value function.

Identify the operation: Generalization uses the available examples rather than treating each one as an isolated fact.

Identify the result: The approximator constructs a broader estimate intended to represent the entire desired function.

Identify the learning setting: This use of examples and desired values makes function approximation an instance of supervised learning within a reinforcement-learning algorithm.

Limited examples serve as evidence for supervised learning, which generalizes from them to construct a broader approximation of the value function.

Key Takeaways

  1. Function approximation turns examples of a desired function into a broader estimate.
  2. Its defining operation is generalization, not simple example storage.
  3. In reinforcement learning, the desired function can be a value function.
  4. Function approximation is an instance of supervised learning used within reinforcement-learning algorithms.
  5. Machine learning, artificial neural networks, pattern recognition, and statistical curve fitting provide possible approximation methods, but practical fit depends on the particular reinforcement-learning algorithm.

Key Takeaways

  • Function approximation uses limited examples to estimate a desired function more broadly.
  • The key operation is generalization from examples, not isolated example storage.
  • The process is supervised learning used within reinforcement-learning algorithms.
  • A desired function in reinforcement learning can be a value function.
  • Methods from several machine-learning and curve-fitting fields may serve as approximators, although their practical fit can differ.