Concepts / Planning with Models

Planning with Models

Models let reinforcement-learning systems simulate experience rather than relying only on directly observed experience.

  • Programming

Simulated Experience

In reinforcement learning, an agent does not always need to obtain every learning experience directly from the environment. A model can imitate the environment well enough to produce possible experience for inspection or planning. This changes the source of the experience: instead of relying only on a direct interaction, the system can ask the model what might happen.

state and selected actionsimulatesPolicyselects actionsModelimitates environmentSimulated experiencepossible outcomes
How does information flow from a policy through a model to create simulated experience instead of using a direct environment interaction?

The central question is not only whether a model produces an outcome. You must also ask what kind of outcome it produces: one possible result, or every possible result together with how likely each result is.

One Simulated Transition

A transition is one simulated step from a starting state after an action is supplied. To interpret this request, focus on two pieces of input: the state and the action. The model uses that state-action information to describe what can happen next.

inputinputproducesproducesStarting statestateActionactionModelsimulates one stepNext statepossible next stateRewardpossible reward
What information enters the model, and how does it produce the next state and reward?

A sample model answers with a single possible way that the step could unfold. A distribution model answers more broadly by generating all possible transitions and attaching their probabilities of occurring. Both model types describe a transition, but they provide different amounts of information about its possible outcomes.

Reading a Transition Request

A planning request supplies a state and an action. Determine the scale of the requested simulation and compare the two possible model outputs.

Identify the input: The input contains a starting state and an action, so the request concerns one transition rather than a complete episode.

Apply a sample model: The sample model returns one possible transition from that state after the action is supplied.

Apply a distribution model: The distribution model returns all possible transitions for that state and action together with their probabilities.

A state plus an action calls for a transition-level interpretation. The model type determines whether the result is one possible transition or the full probability-weighted collection of possible transitions.

Sample and Distribution Outputs

sample modeldistribution modelState and actionsame inputState and actionsame inputOne transitionone possible resultAll transitionswith probabilities
Given a state and action, how does a sample model's single predicted outcome differ from a distribution model's possible outcomes and probabilities?
Model typeInput scaleResult
Sample modelState and action for a transition, or state and policy for an episodeOne possible transition or episode
Distribution modelState and action for a transition, or state and policy for an episodeAll possible transitions or episodes together with their probabilities

What do you think happens?

A model receives a starting state and an action. If it is a distribution model, will it return only one possible next result?

  • Yes, always one result
  • No, all possible transitions with their probabilities
  • It returns an entire episode only
Reveal answer

Answer: No, all possible transitions with their probabilities

A starting state plus an action describes a transition. A distribution model gives the complete set of possible transitions and the probabilities of those transitions.

Generating an Episode

A simulation can continue beyond one transition. When the input consists of a starting state and a policy, the model can extend the simulation into an episode. The policy supplies the actions used while the model supplies possible experience resulting from those actions.

policy selectsmodel simulatesstep continuespolicy selectsadds experienceStarting statestateActionchosen by policyRewardsimulated resultNext statestateNext actionchosen by policyEpisodesequence of experience
How do a starting state and a policy produce a sequence of states, actions, and rewards?

A sample model can produce an entire episode as one possible sequence of experience. In contrast, a distribution model can generate all possible episodes and the probabilities of those episodes. Therefore, a sample-model episode is narrow and path-specific, while a distribution-model result describes the broader set of possible episodes in probabilistic terms.

Interpreting a Policy-Based Simulation

A simulation begins with a state and uses a policy. Determine what an episode-level sample model and an episode-level distribution model produce.

Identify the input: A starting state plus a policy describes an episode-level simulation rather than a single transition.

Use a sample model: The sample model follows one possible sequence of experience generated from the starting state while using the policy.

Use a distribution model: The distribution model describes all possible episodes and gives the probabilities of those episodes.

The same starting state and policy can produce either one possible episode or the full probability-weighted collection of possible episodes, depending on the model type.

Choosing the Right Interpretation

When interpreting a model's output, use two questions in order. First, identify the requested input. A state and action describe a step, so interpret the result as a transition. A state and policy describe an extended simulation, so interpret the result as an episode. Second, identify the model type. A sample model gives one possible result; a distribution model gives every possible result together with its probabilities.

returnsreturnsSample modelone possible resultDistribution modelcomplete outcome setOne transitionor one episodePossible transitionswith probabilities
What kind of result does each model return: one sampled transition or a probability distribution over possible transitions?
  1. Read the input: state and action means transition; state and policy means episode.
  2. Read the model type: sample means one possible result; distribution means all possible results with probabilities.
  3. State the output at the same scale as the input: one transition or one episode for a sample model, and all transitions or all episodes with probabilities for a distribution model.
EASY

A planning system receives a starting state and an action. It uses a sample model. Describe the kind of result the system should interpret as the model's output.

Hints
  • Use the input to identify whether this is a transition or an episode.
  • Use the model type to decide whether the output is one result or a probability-weighted collection.

Common Interpretation Errors

  • Treating a sample model as if it lists every possible outcome.

    A sample model returns one possible transition or episode, not all possible results.

    Fix: Interpret the output as one possible result unless the model is identified as a distribution model.

  • Treating a distribution model as if it returns only one outcome.

    A distribution model returns all possible transitions or episodes together with their probabilities.

    Fix: Look for the complete probability-weighted collection of possible results.

  • Using the wrong simulation scale.

    A state and policy describe an episode-level simulation, while a state and action describe a transition-level simulation.

    Fix: Identify the requested input before interpreting the output.

  • Ignoring the policy when describing an episode.

    The episode-level input consists of a starting state and a policy.

    Fix: Include both the starting state and the policy when explaining how the episode is generated.

Key Takeaways

  1. A model lets a reinforcement-learning system simulate possible experience instead of relying only on directly observed experience.
  2. A state and action describe one simulated transition.
  3. A state and policy describe an episode-level simulation.
  4. A sample model returns one possible transition or episode.
  5. A distribution model returns all possible transitions or episodes together with their probabilities.

Key Takeaways

  • Models imitate an environment well enough to produce possible experience for inspection or planning.
  • The input determines the scale: state plus action means transition, while state plus policy means episode.
  • Sample models provide one possible result.
  • Distribution models provide all possible results together with their probabilities.
  • Correct interpretation requires checking both the input and the model type.