Concepts / Approximate Solution Methods

Approximate Solution Methods

Large state spaces create problems of memory, time, and data, and many encountered states may never have appeared before.

  • Programming

When State Tables Stop Scaling

Imagine preparing a separate answer for every situation an agent might encounter. In a small state space, a table of states and decisions may seem reasonable. In a large state space, however, the challenge is not only storing that table. The agent would also need enough time and data to fill it accurately. Many states it encounters may be new, so remembering exact answers for previously visited states is not enough.

manageableincreasesincreasesincreasesSmall state spaceseparate decisionsMemorylarge tableLarge state spacemany decisionsTimefill the tableDatacover states
What changes when the number of possible states becomes too large for a separate stored decision for every state?

A large state space creates three connected difficulties: memory for storing many decisions, time for preparing them, and data for learning them accurately.

A New State with Familiar Patterns

What do you think happens?

An agent reaches a state it has never seen exactly before, but it has experienced several different states with related characteristics. Can those earlier experiences still help?

  • No. Only an exact match can provide guidance.
  • Yes. Generalization can use experience from different but similar states.
  • Only if every earlier state receives the same decision.
Reveal answer

Answer: Yes. Generalization can use experience from different but similar states.

Generalization uses experience from different but similar states to support sensible decisions in an unseen state. It does not require copying one earlier case mechanically or giving every related state the same answer.

The useful question is not whether the agent has seen the exact current state before. The more useful question is whether earlier encounters can provide guidance for this new state. Generalization turns a limited collection of experience into a basis for decisions across a larger part of the state space.

contributescontributessupportsextends toguidesObserved state Arelated characteristicsPrior experienceexamplesGeneralizationguidanceUnseen statenew caseDecisionsupported choiceObserved state Brelated characteristics
How can an agent use what it learned from similar states when it encounters a state it has never seen before?

Using related experience

An agent has encountered several situations with related characteristics. It later encounters a different situation that is not an exact match for any earlier one. How can the earlier encounters help?

Collect experience: The agent has examples from several previously encountered states.

Identify relevance: The new state is different from the earlier states, but it is not completely unrelated to them.

Generalize: The agent uses the earlier examples to form guidance for the new state rather than requiring an exact remembered answer.

Support a decision: The generalized guidance can influence the decision for the new state.

Experience from similar states can support a decision in an unseen state. Generalization does not mean mechanically copying one remembered case or assigning the same answer to every related state.

Function Approximation

Function approximation is a standard way to implement generalization. It begins with examples from a desired function, such as a value function, and uses those examples to construct an approximation over a much larger set of cases. In reinforcement learning, this lets what was learned from particular encounters extend beyond those exact examples.

describesinformsconstructssupportsState informationobserved caseFunction examplesobserved casesFunction approximatorgeneralized constructionEstimated valueadditional caseDecisionadditional case
How does a function approximator transform observed state information into an estimated value or decision?
Desired functionApproximation
The function the learner is trying to representThe generalized construction produced from available examples
May be described as a value functionIs intended to cover more cases than the examples directly provide

The Supervised Learning Connection

Function approximation is an instance of supervised learning. This means that a set of examples can be used with a supervised learning method to play the role of a function approximator inside a reinforcement learning algorithm. The connection allows reinforcement learning to draw on methods from machine learning and related fields.

provided toproducesextends acrossObserved examplesdesired functionSupervised learninglearning methodFunction approximatorgeneralized functionLarger set of casesadditional states
How do observed examples become part of a supervised learning method used in reinforcement learning?

When analyzing an approximate solution method, keep the chain clear: observed examples provide information about a desired function; a supervised learning method uses those examples; the resulting approximation extends guidance to more cases than the examples directly contain.

Mistakes About Generalization

  • Assuming an exact lookup table is sufficient for a large state space.

    A large state space creates memory, time, and data difficulties, and many states encountered by the agent may be new.

    Fix: Use generalization so experience from different but similar states can support decisions for unseen states.

  • Treating generalization as copying one earlier decision.

    Generalization forms an approximation from examples; it does not mean mechanically copying one remembered case.

    Fix: Understand the earlier examples as evidence used to construct guidance over additional cases.

  • Confusing the desired function with its approximation.

    The desired function is what the learner is trying to represent, while the approximation is the construction produced from available examples.

    Fix: Keep the goal function and the learned approximation conceptually separate.

  • Treating function approximation as unrelated to supervised learning.

    Function approximation is an instance of supervised learning.

    Fix: Connect the examples, the supervised learning method, and the resulting approximation.

Check Your Understanding

MEDIUM

Explain why an agent operating in a large state space cannot rely only on a separate remembered answer for every state it has already visited. Then describe how function approximation changes the role of the agent's experience.

Hints
  • Mention memory, time, and data requirements.
  • Explain why an encountered state may be new.
  • Distinguish examples of a desired function from the approximation built from those examples.
MEDIUM

A new state is not an exact match for any previously observed state, but it has related characteristics. In a short paragraph, explain why earlier examples may still provide useful guidance without requiring the agent to assign the new state exactly the same answer as an earlier state.

Hints
  • Use the term generalization.
  • Explain the role of an approximation.
  • Avoid describing generalization as simple copying.

Key Takeaways

  1. Large state spaces make separate state-by-state decisions difficult to store, prepare, and learn accurately. Generalization allows experience from different but similar states to guide decisions for unseen states. Function approximation constructs an approximation of a desired function from available examples and extends it over a larger set of cases. Because function approximation is an instance of supervised learning, reinforcement learning can use methods from machine learning and related fields.

Key Takeaways

  • Large state spaces create problems of memory, time, and data.
  • Many states an agent encounters may never have appeared in its previous experience.
  • Generalization uses experience from different but similar states to support decisions in unseen states.
  • Function approximation builds a generalized construction from examples of a desired function.
  • Function approximation is an instance of supervised learning.