Environment Models in Reinforcement Learning
Cognitive maps and environment models are representations of environments, described respectively in psychological and reinforcement learning terms.
From Experience to Planning
Imagine an animal encountering an environment before it has a reason to travel toward any particular location. Later, something changes unexpectedly, and the animal must decide how to behave. The important question is whether it can use an internal representation of how the environment is organized, rather than relying only on previously learned state-action associations. Psychology calls this kind of mental representation a cognitive map. Reinforcement learning uses the related idea of an environment model: a computational representation that mimics how the environment behaves.
Cognitive Maps and Environment Models
A cognitive map is a mental representation used to understand and navigate an environment. An environment model is a computational representation that an agent can use to predict how the environment behaves.
These terms come from different fields and should not be treated as names for identical physical mechanisms. Psychology uses cognitive map to describe a mental representation of an environment. Reinforcement learning uses environment model for a representation that mimics environmental behavior. The connection is a shared pattern: experience can produce a representation, and that representation can later support planning rather than merely storing a record of past actions.
Representation Beyond Reward
Learning an environment representation is different from learning only which actions have previously produced rewards. The source identifies a shared learning property of cognitive maps and environment models: they can be learned without relying on reward signals. Their later purpose is planning behavior. This means the representation concerns how the environment is organized or how it responds, not simply a list of rewarded state-action associations.
| Learning emphasis | What is represented | Later use |
|---|---|---|
| Environment representation | How the environment is organized or behaves | Planning behavior |
| State-action associations | Connections between situations, actions, and learned behavior | Guiding behavior through learned associations |
Planning After Change
A Changed Route
An agent has learned a representation of an environment. Later, an unexpected change makes its previously useful route unsuitable. How can the representation help?
Recognize the changed situation: The agent does not have to treat every possible response as completely unknown. It can use its representation as information about the environment.
Consider alternatives: The agent can use the representation to think about how different actions may affect what happens next.
Choose a new action: Planning with the representation can guide behavior after the unexpected change instead of requiring reliance only on an old state-action association.
The learned representation supports planning when the environment changes unexpectedly.
What a Model Predicts
In reinforcement learning, a model is anything the agent can use to predict how the environment will react to an action. The prediction starts with a state and an action and concerns the next state and reward.
Without a model, an agent must wait for the environment to respond whenever it wants to consider an action. With a model, the agent can predict that response. The model therefore connects a possible state-action situation to its possible environmental consequences: the next state and the reward.
Distribution and Sample Models
Distribution models and sample models differ in how much of the possible response they return. A distribution model exposes the complete set of possibilities and their probabilities. A sample model exposes one probability-based possibility at a time. The distribution model is therefore more expressive because it describes the full response distribution, while the sample model gives one possible response selected according to that distribution.
| Model type | What it returns | Expressiveness | Practical characterization |
|---|---|---|---|
| Distribution model | The complete set of possibilities and their probabilities | More expressive | Describes the full response distribution |
| Sample model | One probability-based possibility at a time | Less expressive than a complete distribution | Can be easier to obtain |
Choosing the Model View
One Input, Two Model Outputs
Suppose an agent asks a model about one state and one action. What would the two model forms provide?
Distribution model: It provides the complete set of possible next states and rewards together with their probabilities.
Sample model: It provides one probability-based possible next state and reward at a time.
Interpret the difference: The distribution model gives more information about all possible responses. The sample model gives a single possible response and can be easier to obtain.
The two models can describe the same environmental response pattern at different levels of detail.
Common Misunderstandings
Treating a cognitive map and an environment model as identical terms.
The source says the ideas are related representations described in different fields, not identical mechanisms.
Fix:
Describe their shared pattern while keeping the terms distinct: a cognitive map is a psychological mental representation, whereas an environment model is a reinforcement-learning computational representation.Assuming that an environment representation must be learned from reward.
The source identifies learning without relying on reward signals as a shared property of these representations.
Fix:
Separate learning the environment representation from learning directly from reward signals.Defining a model as a record of past actions.
A model is defined by its use in predicting how the environment reacts to an action.
Fix:
Ask whether the representation predicts the next state and reward from a state and an action.Confusing a distribution model with a sample model.
That describes a distribution model. A sample model returns one probability-based possibility at a time.
Fix:
Check whether the output is the complete set of possibilities or one sampled possibility.
Check Your Understanding
An agent has a representation that, given a state and an action, provides one probability-based possible next state and reward. Identify the model type and explain how its output differs from that of a distribution model.
Hints
- Look for whether the model returns one possibility or the complete set of possibilities.
- Mention probabilities in both descriptions.
What do you think happens?
An environment changes unexpectedly after an agent has learned a representation of it. Should the agent be limited to repeating its old state-action associations?
Reveal answer
Answer: No, because the representation can support planning about alternatives.
The source describes cognitive maps and environment models as representations that can later guide planning, including when behavior must respond to an unexpected change. This does not mean the model eliminates the environment or that reward is required to create the representation.
Key Takeaways
- A cognitive map is a mental representation of an environment used to understand and navigate it.
- An environment model is a computational representation that predicts how the environment reacts to an action.
- Both representations can be learned without relying on reward signals and can later support planning.
- A model's prediction begins with a state and an action and concerns the next state and reward.
- A distribution model returns all possibilities and their probabilities, whereas a sample model returns one probability-based possibility at a time.
Key Takeaways
- Cognitive maps and environment models are related but field-specific terms for representations of environments.
- Their important role is to support planning, especially when behavior must adapt to an unexpected change.
- Learning an environment representation is not the same as learning only state-action associations from reward signals.
- A reinforcement-learning model predicts the next state and reward from a state and an action.
- Distribution models describe complete possibilities and probabilities; sample models return one probability-based possibility at a time.