Concepts / Models and Planning Fundamentals

Models and Planning Fundamentals

Models let agents predict environment responses instead of treating every possible response as unknown.

  • Programming

From Reaction to Prediction

An agent normally chooses an action and waits for the environment to respond. A model changes this process: instead of treating every possible response as unknown, the agent can use a model to predict how the environment will react. In reinforcement learning, a model is anything the agent can use to predict the environment's response to an action.

A model is a prediction tool for environment responses. It is not the action itself and it is not the agent's chosen policy; it describes what may happen after a state and action.

The Prediction Path

The model's prediction starts with two inputs: the current state and a chosen action. From those inputs, the prediction concerns two parts of the environment response: the next state and the reward. The next state describes where the environment goes next, while the reward is the response value associated with that transition. The model therefore connects a possible choice with its possible consequences.

inputinputpredictspredictsCurrent stateModelpredicts responseNext stateChosen actionReward
Given a current state and a chosen action, how does the model produce the predicted next state and reward?

A Hypothetical Environment Response

Suppose an agent is in a state called Corridor and considers the action Move forward. What does the model need to predict?

Identify the inputs: The model receives the state Corridor and the action Move forward.

Identify the predicted response: The model predicts a next state and a reward associated with that state-action choice.

Separate prediction from experience: This prediction lets the agent consider the response without treating the environment's actual response as the only available source of information.

The model predicts the environment response from the pair of inputs: current state plus chosen action. That response concerns the next state and reward.

Two Ways to Return Possibilities

Models differ in how much of the possible environment response they return. A distribution model describes the complete set of possibilities and gives the probabilities associated with them. A sample model returns one possibility at a time, with that possibility selected according to the probabilities.

includesincludesreturnsDistribution modelcomplete possibilities andprobabilitiesOutcome AprobabilitySelected outcomeone possibilityOutcome BprobabilitySample modelone probability-basedpossibility
What is the difference between a model that returns all possible outcomes with their probabilities and one that returns only a single sampled outcome?
Model typeWhat it returnsInformation represented
Distribution modelThe complete set of possibilities and their probabilitiesThe full possible response distribution
Sample modelOne probability-based possibility at a timeA single outcome from the response distribution

The distinction concerns how much of the possible response the model exposes.

One Response or the Whole Set

Imagine that a state-action pair could lead to two possible responses. How would the two model types represent that situation?

Distribution representation: A distribution model returns both possible responses together with their probabilities.

Sample representation: A sample model returns only one of those possible responses for the particular sample it produces.

Compare the information: The distribution representation preserves the complete set of possibilities, while the sample representation exposes one probability-based possibility at a time.

Both model types describe the same kind of environment response, but they expose different amounts of that response in each use.

Planning Without Immediate Interaction

Planning uses a model to consider what may happen after an action. The agent can start from a state, consider an action, ask the model for the predicted response, and use that prediction while deciding what to do next. This means the agent does not need to wait for the environment to respond every time it wants to consider an action.

consideruse modelinform decisionnext planning stateCurrent stateCandidate actionPredicted responsenext state and rewardChosen action
How does an agent use predicted responses to simulate possible actions and choose what to do next?

Expressiveness and Practicality

Distribution models are more expressive because they expose the complete set of possible responses and the probabilities of those responses. A sample model exposes only one probability-based possibility at a time, so less of the full response distribution is visible in each result. However, sample models can be easier to obtain. The key tradeoff is therefore between preserving the full distribution and returning individual samples.

preservesreturnsCompletedistributionall possibilities andprobabilitiesMore responseinformationOne response at atimeSingle sampleone possibility
What information is preserved by a full outcome distribution, and how does a sample model differ in what it returns?

The model form used in dynamic-programming estimates of an MDP's dynamics is a distribution model. The blackjack example described in the source material uses a sample model. These are applications of the two distinct ways of representing environment responses.

Common Misunderstandings

  • Thinking that a model predicts only the next state.

    The prediction from a state and action concerns both the next state and the reward.

    Fix: When describing a model response, name both predicted parts: next state and reward.

  • Assuming that a sample model returns the complete set of possible outcomes.

    A sample model exposes one probability-based possibility at a time.

    Fix: Reserve the complete set of possibilities and their probabilities for a distribution model.

  • Assuming that planning always requires an immediate environment response.

    An agent can use a model to predict the response while considering an action.

    Fix: Separate model-based prediction from direct interaction with the environment.

  • Treating a sample model as more informative than a distribution model.

    A distribution model exposes the complete possibilities and probabilities, while a sample model returns one possibility at a time.

    Fix: Remember that distribution models are more expressive, even though sample models can be easier to obtain.

Check Your Understanding

EASY

An agent is in state S and considers action A. Explain what a model predicts, then describe how a distribution model and a sample model would return that prediction.

Hints
  • Start with the two inputs: state and action.
  • Name both parts of the predicted environment response.
  • For the distribution model, describe all possibilities and their probabilities.
  • For the sample model, describe the single probability-based possibility returned at a time.

What do you think happens?

An agent wants to consider an action but does not want to wait for the environment to respond immediately. Which resource can it use?

  • A model that predicts the environment response
  • Only a direct environment response
  • A reward without a state and action
  • A model that returns no possible response
Reveal answer

Answer: A model that predicts the environment response

An agent can use a model to predict the response to an action instead of waiting for the environment every time it considers an action.

Essential Takeaways

  1. In reinforcement learning, a model is anything an agent can use to predict how the environment will react to an action.
  2. The prediction begins with a state and an action and concerns the next state and reward.
  3. A distribution model returns the complete set of possibilities and their probabilities.
  4. A sample model returns one probability-based possibility at a time.
  5. Distribution models are more expressive because they preserve the complete response distribution, while sample models can be easier to obtain.

Key Takeaways

  • A reinforcement-learning model predicts an environment response from a state and an action.
  • The predicted response concerns the next state and reward.
  • Distribution models expose all possible responses and their probabilities.
  • Sample models expose one probability-based response at a time.
  • Planning can use model predictions so the agent does not need to wait for a direct environment response every time it considers an action.