Models and Planning Fundamentals
Models let agents predict environment responses instead of treating every possible response as unknown.
From Reaction to Prediction
An agent normally chooses an action and waits for the environment to respond. A model changes this process: instead of treating every possible response as unknown, the agent can use a model to predict how the environment will react. In reinforcement learning, a model is anything the agent can use to predict the environment's response to an action.
A model is a prediction tool for environment responses. It is not the action itself and it is not the agent's chosen policy; it describes what may happen after a state and action.
The Prediction Path
The model's prediction starts with two inputs: the current state and a chosen action. From those inputs, the prediction concerns two parts of the environment response: the next state and the reward. The next state describes where the environment goes next, while the reward is the response value associated with that transition. The model therefore connects a possible choice with its possible consequences.
A Hypothetical Environment Response
Suppose an agent is in a state called Corridor and considers the action Move forward. What does the model need to predict?
Identify the inputs: The model receives the state Corridor and the action Move forward.
Identify the predicted response: The model predicts a next state and a reward associated with that state-action choice.
Separate prediction from experience: This prediction lets the agent consider the response without treating the environment's actual response as the only available source of information.
The model predicts the environment response from the pair of inputs: current state plus chosen action. That response concerns the next state and reward.
Two Ways to Return Possibilities
Models differ in how much of the possible environment response they return. A distribution model describes the complete set of possibilities and gives the probabilities associated with them. A sample model returns one possibility at a time, with that possibility selected according to the probabilities.
| Model type | What it returns | Information represented |
|---|---|---|
| Distribution model | The complete set of possibilities and their probabilities | The full possible response distribution |
| Sample model | One probability-based possibility at a time | A single outcome from the response distribution |
The distinction concerns how much of the possible response the model exposes.
One Response or the Whole Set
Imagine that a state-action pair could lead to two possible responses. How would the two model types represent that situation?
Distribution representation: A distribution model returns both possible responses together with their probabilities.
Sample representation: A sample model returns only one of those possible responses for the particular sample it produces.
Compare the information: The distribution representation preserves the complete set of possibilities, while the sample representation exposes one probability-based possibility at a time.
Both model types describe the same kind of environment response, but they expose different amounts of that response in each use.
Planning Without Immediate Interaction
Planning uses a model to consider what may happen after an action. The agent can start from a state, consider an action, ask the model for the predicted response, and use that prediction while deciding what to do next. This means the agent does not need to wait for the environment to respond every time it wants to consider an action.
Expressiveness and Practicality
Distribution models are more expressive because they expose the complete set of possible responses and the probabilities of those responses. A sample model exposes only one probability-based possibility at a time, so less of the full response distribution is visible in each result. However, sample models can be easier to obtain. The key tradeoff is therefore between preserving the full distribution and returning individual samples.
The model form used in dynamic-programming estimates of an MDP's dynamics is a distribution model. The blackjack example described in the source material uses a sample model. These are applications of the two distinct ways of representing environment responses.
Common Misunderstandings
Thinking that a model predicts only the next state.
The prediction from a state and action concerns both the next state and the reward.
Fix:
When describing a model response, name both predicted parts: next state and reward.Assuming that a sample model returns the complete set of possible outcomes.
A sample model exposes one probability-based possibility at a time.
Fix:
Reserve the complete set of possibilities and their probabilities for a distribution model.Assuming that planning always requires an immediate environment response.
An agent can use a model to predict the response while considering an action.
Fix:
Separate model-based prediction from direct interaction with the environment.Treating a sample model as more informative than a distribution model.
A distribution model exposes the complete possibilities and probabilities, while a sample model returns one possibility at a time.
Fix:
Remember that distribution models are more expressive, even though sample models can be easier to obtain.
Check Your Understanding
An agent is in state S and considers action A. Explain what a model predicts, then describe how a distribution model and a sample model would return that prediction.
Hints
- Start with the two inputs: state and action.
- Name both parts of the predicted environment response.
- For the distribution model, describe all possibilities and their probabilities.
- For the sample model, describe the single probability-based possibility returned at a time.
What do you think happens?
An agent wants to consider an action but does not want to wait for the environment to respond immediately. Which resource can it use?
Reveal answer
Answer: A model that predicts the environment response
An agent can use a model to predict the response to an action instead of waiting for the environment every time it considers an action.
Essential Takeaways
- In reinforcement learning, a model is anything an agent can use to predict how the environment will react to an action.
- The prediction begins with a state and an action and concerns the next state and reward.
- A distribution model returns the complete set of possibilities and their probabilities.
- A sample model returns one probability-based possibility at a time.
- Distribution models are more expressive because they preserve the complete response distribution, while sample models can be easier to obtain.
Key Takeaways
- A reinforcement-learning model predicts an environment response from a state and an action.
- The predicted response concerns the next state and reward.
- Distribution models expose all possible responses and their probabilities.
- Sample models expose one probability-based response at a time.
- Planning can use model predictions so the agent does not need to wait for a direct environment response every time it considers an action.