Understanding Models and Planning in Reinforcement Learning
Models let agents predict environment responses instead of treating every possible response as unknown.
When the Environment Is Not the Only Source of Answers
An agent often has to consider what might happen if it takes an action. Without a model, it must treat the environment's response as unknown until the action is actually carried out. With a model, the agent can predict the response and consider actions without waiting for the environment to respond every time.
In reinforcement learning, a model is anything the agent can use to predict how the environment will react to an action.
The Model's Prediction
A model prediction begins with two pieces of information: the current state and an action the agent is considering. The prediction concerns two parts of the environment's response: the reward and the next state. In other words, the model helps answer this question: if the agent is in this state and takes this action, what reward and next state should it expect?
Tracing One Model Query
An agent is in the generated state named Junction A and considers the generated action Move east. What information does the model need to predict?
Start with the state: The model receives Junction A as the state in which the agent is currently situated.
Add the action: The model also receives Move east, the action the agent is considering.
Predict the response: The model predicts the reward associated with this response and the next state reached after the action.
The model query is based on a state and an action, and its predicted response consists of a reward and a next state.
Two Ways to Represent Outcomes
The important difference between distribution models and sample models is how much of the possible response the model returns. A distribution model describes the complete set of possibilities and their probabilities. A sample model returns one possibility selected according to those probabilities.
| Feature | Distribution model | Sample model |
|---|---|---|
| Returned information | The complete set of possibilities | One possibility at a time |
| Probability information | Includes the probabilities of the possibilities | Uses the probabilities to select one possibility |
| Expressiveness | More expressive because it exposes the full set of possibilities and probabilities | Less expressive for a single query because it exposes one possibility |
| Source-grounded application | Used in dynamic-programming estimates of an MDP's dynamics | Used in the blackjack example described in the source material |
Why Representation Changes Planning
Planning means considering possible actions by using predicted environment responses. Instead of treating every response as unknown, the agent can query its model, examine the predicted reward and next state, and use that information while deciding what to do. The model therefore gives the agent a way to consider actions without requiring an immediate response from the environment for every consideration.
Planning with Different Model Outputs
Suppose an agent is considering an action from a generated state called Room 1. How would planning differ depending on the model type?
Use a distribution model: The agent receives the complete set of possible responses to the state and action, together with their probabilities. It can therefore consider the full range of predicted outcomes.
Use a sample model: The agent receives one probability-based possibility for the state and action. It can use that possible response as one simulated result.
Compare the meaning: The distribution model exposes more information about the possible response, while the sample model provides one outcome at a time.
Both model types let the agent predict an environment response, but they provide different amounts of information about the possible outcomes.
Mistakes About Models
Treating a model as a prediction of only the next state
The model prediction concerns both the next state and the reward.
Fix:
When describing a model response, name both predicted parts: reward and next state.Assuming that a model is the environment's actual response
A model lets the agent predict how the environment will react without waiting for the environment to respond every time it considers an action.
Fix:
Treat the model as a way to predict an environment response for planning.Confusing one sampled outcome with the complete distribution
A sample model exposes one probability-based possibility at a time; the distinction does not say that no other possibilities exist.
Fix:
Use distribution model for the complete set and probabilities, and sample model for one selected possibility.Calling distribution and sample models different kinds of predicted input
Both forms begin with a state and an action. They differ in how much of the possible response they return.
Fix:
Keep the input fixed in your explanation and compare the returned outcome information.
Practice the Distinction
An agent considers an action from a current state. One model describes every possible reward and next-state outcome along with the probability of each. Another model returns one probability-based reward and next-state outcome. Identify which model is the distribution model and which is the sample model, then explain why.
Hints
- Look at whether the model returns the complete set of possibilities or one possibility.
- Probability information is explicitly represented in the distribution model.
- The sample model returns one possibility selected according to probabilities.
- A correct answer identifies the first model as a distribution model because it exposes every possible outcome and its probability. The second is a sample model because it exposes one probability-based possibility at a time.
Key Takeaways
- A reinforcement learning model is anything an agent can use to predict how the environment will react to an action.
- The prediction starts with a state and an action and concerns a reward and a next state.
- A distribution model represents the complete set of possible responses and their probabilities.
- A sample model returns one probability-based possibility at a time.
- Models support planning because an agent can consider predicted responses without waiting for the environment to respond every time.
Key Takeaways
- A model predicts the environment's response to an action from a given state.
- The predicted response concerns both reward and next state.
- Distribution models expose all possible outcomes and their probabilities.
- Sample models expose one probability-based outcome at a time.
- Using a model allows an agent to consider actions and plan without treating every response as unknown.