Planning with Models
Models let reinforcement-learning systems simulate experience rather than relying only on directly observed experience.
Simulated Experience
In reinforcement learning, an agent does not always need to obtain every learning experience directly from the environment. A model can imitate the environment well enough to produce possible experience for inspection or planning. This changes the source of the experience: instead of relying only on a direct interaction, the system can ask the model what might happen.
The central question is not only whether a model produces an outcome. You must also ask what kind of outcome it produces: one possible result, or every possible result together with how likely each result is.
One Simulated Transition
A transition is one simulated step from a starting state after an action is supplied. To interpret this request, focus on two pieces of input: the state and the action. The model uses that state-action information to describe what can happen next.
A sample model answers with a single possible way that the step could unfold. A distribution model answers more broadly by generating all possible transitions and attaching their probabilities of occurring. Both model types describe a transition, but they provide different amounts of information about its possible outcomes.
Reading a Transition Request
A planning request supplies a state and an action. Determine the scale of the requested simulation and compare the two possible model outputs.
Identify the input: The input contains a starting state and an action, so the request concerns one transition rather than a complete episode.
Apply a sample model: The sample model returns one possible transition from that state after the action is supplied.
Apply a distribution model: The distribution model returns all possible transitions for that state and action together with their probabilities.
A state plus an action calls for a transition-level interpretation. The model type determines whether the result is one possible transition or the full probability-weighted collection of possible transitions.
Sample and Distribution Outputs
| Model type | Input scale | Result |
|---|---|---|
| Sample model | State and action for a transition, or state and policy for an episode | One possible transition or episode |
| Distribution model | State and action for a transition, or state and policy for an episode | All possible transitions or episodes together with their probabilities |
What do you think happens?
A model receives a starting state and an action. If it is a distribution model, will it return only one possible next result?
Reveal answer
Answer: No, all possible transitions with their probabilities
A starting state plus an action describes a transition. A distribution model gives the complete set of possible transitions and the probabilities of those transitions.
Generating an Episode
A simulation can continue beyond one transition. When the input consists of a starting state and a policy, the model can extend the simulation into an episode. The policy supplies the actions used while the model supplies possible experience resulting from those actions.
A sample model can produce an entire episode as one possible sequence of experience. In contrast, a distribution model can generate all possible episodes and the probabilities of those episodes. Therefore, a sample-model episode is narrow and path-specific, while a distribution-model result describes the broader set of possible episodes in probabilistic terms.
Interpreting a Policy-Based Simulation
A simulation begins with a state and uses a policy. Determine what an episode-level sample model and an episode-level distribution model produce.
Identify the input: A starting state plus a policy describes an episode-level simulation rather than a single transition.
Use a sample model: The sample model follows one possible sequence of experience generated from the starting state while using the policy.
Use a distribution model: The distribution model describes all possible episodes and gives the probabilities of those episodes.
The same starting state and policy can produce either one possible episode or the full probability-weighted collection of possible episodes, depending on the model type.
Choosing the Right Interpretation
When interpreting a model's output, use two questions in order. First, identify the requested input. A state and action describe a step, so interpret the result as a transition. A state and policy describe an extended simulation, so interpret the result as an episode. Second, identify the model type. A sample model gives one possible result; a distribution model gives every possible result together with its probabilities.
- Read the input: state and action means transition; state and policy means episode.
- Read the model type: sample means one possible result; distribution means all possible results with probabilities.
- State the output at the same scale as the input: one transition or one episode for a sample model, and all transitions or all episodes with probabilities for a distribution model.
A planning system receives a starting state and an action. It uses a sample model. Describe the kind of result the system should interpret as the model's output.
Hints
- Use the input to identify whether this is a transition or an episode.
- Use the model type to decide whether the output is one result or a probability-weighted collection.
Common Interpretation Errors
Treating a sample model as if it lists every possible outcome.
A sample model returns one possible transition or episode, not all possible results.
Fix:
Interpret the output as one possible result unless the model is identified as a distribution model.Treating a distribution model as if it returns only one outcome.
A distribution model returns all possible transitions or episodes together with their probabilities.
Fix:
Look for the complete probability-weighted collection of possible results.Using the wrong simulation scale.
A state and policy describe an episode-level simulation, while a state and action describe a transition-level simulation.
Fix:
Identify the requested input before interpreting the output.Ignoring the policy when describing an episode.
The episode-level input consists of a starting state and a policy.
Fix:
Include both the starting state and the policy when explaining how the episode is generated.
Key Takeaways
- A model lets a reinforcement-learning system simulate possible experience instead of relying only on directly observed experience.
- A state and action describe one simulated transition.
- A state and policy describe an episode-level simulation.
- A sample model returns one possible transition or episode.
- A distribution model returns all possible transitions or episodes together with their probabilities.
Key Takeaways
- Models imitate an environment well enough to produce possible experience for inspection or planning.
- The input determines the scale: state plus action means transition, while state plus policy means episode.
- Sample models provide one possible result.
- Distribution models provide all possible results together with their probabilities.
- Correct interpretation requires checking both the input and the model type.