Models of the Environment
The key distinction is whether the model exposes the full probability description or only one sampled outcome.
Looking Beyond the Current State
An agent does not need a model merely to name the situation it is currently in. It needs a model when it wants to anticipate what the environment may do after it chooses an action. The central question is: given a state and an action, what next state and next reward might result?
A model of the environment gives an agent a way to infer how the environment may respond. The model receives a state-action pair and predicts an environment response involving both a next state and a next reward. This is a prediction about what could happen; the model itself has not caused the real environment to move.
For example, an agent may be in state S1 and consider action A1. A model can receive that pair and produce a prediction that the environment may move to state S2 and provide reward R1. The real environment has not moved because of this prediction. The model has supplied an inference about a possible response.
Two Ways to Report a Response
The main distinction between model types is how much of the possible future they expose at one time. A distribution model lays out the complete probability description of possible outcomes. A sample model gives one outcome selected according to those probabilities.
| Model type | What it returns | What the output represents |
|---|---|---|
| Distribution model | All possible responses together with their probabilities | The complete probability description |
| Sample model | One next-state and reward outcome | One outcome drawn according to the probabilities |
The Dozen-Dice Test Case
Consider a process that rolls a dozen dice and records the sum. The underlying process is the same whether an agent uses a distribution model or a sample model. What changes is the model's report.
Reporting the Sum of a Dozen Dice
Identify what a distribution model and a sample model return for a process that rolls a dozen dice and records the sum.
Distribution model: It returns all possible sums together with the probabilities of those sums. This exposes the complete probability description.
Sample model: It returns one sum selected according to the probabilities. It does not expose all possible sums in that single response.
Compare the underlying process: The dice process and its probability distribution have not changed. Only the amount of information reported by the model has changed.
The distribution model reports the full set of possible sums and their probabilities; the sample model reports one sampled sum.
What do you think happens?
A process rolls a dozen dice and records the sum. Which output best describes a sample model?
Reveal answer
Answer: One sum selected according to the probabilities
A sample model supplies one outcome drawn from the underlying probability distribution. A distribution model would expose all possible sums and their probabilities.
Capability and Construction Trade-Offs
A distribution model contains enough information to generate samples. Because it includes the complete probability description, it can be used to produce the kind of single outcome returned by a sample model. In this sense, a distribution model is more capable.
Greater capability does not mean that a distribution model is always easier to obtain. In many applications, constructing a sample model may be easier. The dozen-dice example illustrates why: simulating the rolls and returning their sum is easy, while determining every possible sum and its probability can be harder and more error-prone.
| Question | Distribution model | Sample model |
|---|---|---|
| How much does it reveal? | The complete probability description | One outcome |
| Can it support a single sampled outcome? | Yes | It already returns one |
| Why might it be difficult or easy to obtain? | Describing every possible outcome and its probability can be harder and more error-prone | Simulating a response and returning one result may be easier |
Planning Before Experience
Planning means considering possible future situations before those situations are actually experienced. A model makes this possible by predicting how the environment may respond to a state and an action. An agent can use those predicted responses to consider consequences before waiting for the real environment to respond.
The model's prediction is not the same thing as an actual transition in the environment. It supports reasoning about possible consequences. This is why models are connected to planning: they allow an agent to consider possible future situations before those situations are actually encountered.
Model-Based and Model-Free Methods
Model-based methods use a model of the environment and planning. Their decisions involve considering possible future situations. Model-free methods do not use a model of the environment and are explicitly trial-and-error learners. They learn through interaction rather than first using a model to reason about possible futures.
The distinction is about whether the method uses a model to reason about possible futures. Model-based methods plan with model-supported predictions. Model-free methods are described as almost the opposite of planning because they learn through interaction and trial and error without first using a model to reason about those possible futures.
Mistakes in Reading Model Outputs
Treating a sample model as if it returned the complete probability distribution.
A sample model returns one outcome selected according to the probabilities.
Fix:
Reserve the complete probability description for a distribution model.Assuming that a distribution model and a sample model describe different underlying processes.
The underlying process has not changed. The model's report has changed.
Fix:
Separate the environment's probability distribution from the amount of that distribution exposed by the model.Thinking that a model causes the real environment to transition.
The model supplies an inference about what could happen; it has not caused the real environment to move.
Fix:
Describe the model output as a predicted or possible response.Calling a method model-based merely because it interacts with an environment.
The model-based and model-free distinction concerns whether the method uses a model and planning.
Fix:
Identify whether the method uses model-supported predictions of possible future situations or learns through trial and error without a model.
Check Your Understanding
An agent gives a state and an action to a model. The model returns one possible next state and one next reward, rather than listing every possible response and its probability. Identify the model type and explain how this output could still support planning.
Hints
- Focus on whether the model returns the full probability description or one outcome.
- Planning uses model predictions to consider possible future situations before they are experienced.
A dozen-dice process is used to record a sum. Write a short comparison of the output from a distribution model and the output from a sample model. Include what remains unchanged between the two cases.
Hints
- A distribution model exposes all possible sums and their probabilities.
- A sample model returns one sum.
- The underlying process and probability distribution have not changed.
Key Takeaways
- A model of the environment helps an agent infer how the environment may respond after a state and an action.
- The predicted response concerns both a next state and a next reward.
- A distribution model exposes all possible outcomes together with their probabilities.
- A sample model returns one outcome selected according to the probabilities.
- Distribution models are more capable, while sample models may be easier to construct.
- Model-based methods use models and planning; model-free methods learn through trial and error without a model.
Key Takeaways
- A model predicts how an environment may respond to a state and an action, including a possible next state and next reward.
- A distribution model returns the complete probability description of possible responses.
- A sample model returns one response drawn according to those probabilities.
- The dozen-dice example shows that the underlying process can stay the same while the model reports either all possible sums with probabilities or one sum.
- Model-based methods plan with models, whereas model-free methods learn through trial and error without a model.