Cognitive Maps in Reinforcement Learning
Latent learning separates the occurrence of learning from its immediate expression in behavior.
Learning Before Action
An agent can learn about an environment before it has a reason to act toward a particular goal. During that period, its behavior may not clearly reveal what it has learned. If a goal later becomes important, the earlier learning can become visible through planning. This separation between learning and its immediate expression in behavior is called latent learning.
What do you think happens?
An animal explores an environment without receiving a reward or penalty. Later, reaching a particular location becomes important. What is the strongest interpretation if the animal takes a different route to that location rather than merely repeating a practiced sequence?
Reveal answer
Answer: The animal may have learned an environment representation that now supports flexible planning.
Tolman's interpretation was that learning can occur before it becomes visible in goal-directed behavior. A different route matters because it suggests the use of relationships represented in an environment model rather than simple repetition of an earlier reinforced response.
Tolman's Cognitive Map
A cognitive map is a mental representation used to understand and navigate an environment. In Tolman's account, it is a learned model of an environment or task space. The representation can capture relationships in the environment so that behavior is not limited to repeating one previously reinforced sequence.
The cognitive-map perspective begins with experience. By encountering an environment, an animal can learn information about the environment even when no reward or penalty is guiding that original learning. The resulting representation can later help the animal decide how to reach a relevant goal. Rewards or penalties can make an already learned model relevant to planning without being required for the original learning.
From Model to Route
A Newly Relevant Goal
An animal has explored an environment without receiving a reward or penalty. Later, reaching a particular location becomes important. How can its earlier learning support behavior now?
Explore: The animal encounters the environment and gathers information about it. Its behavior during this period may not provide an obvious sign that learning has occurred.
Represent: The experience contributes to a learned model of the environment or task space. This model represents relationships that can later be useful for navigation.
Motivate: A goal becomes important. The goal does not have to be the signal that originally created the representation; it makes the already learned model relevant to current behavior.
Plan: The animal can use the represented relationships to select a route to the goal. If it uses a different route rather than repeating one practiced sequence, the behavior illustrates flexible planning.
Latent learning becomes visible as goal-directed behavior supported by an environment model.
In reinforcement learning terms, this kind of environment model is the computational counterpart of the representational idea described by the cognitive map. A model-based approach can use a learned representation of how the environment behaves to support planning. The important contrast is with an explanation based only on state-action associations: the model allows the agent to use relationships in the environment when selecting behavior.
Representation Versus Reward Association
| Learning emphasis | What is learned | Later use described by the source |
|---|---|---|
| Environment representation | Information about the environment, including relationships in the task space | Planning behavior and selecting a route, including a different route to a goal |
| Reward-based association | State-action associations connected with earlier reinforcement | Explaining behavior through previously reinforced responses |
The cognitive-map perspective does not require behavior to be understood only as a collection of state-action associations. The source describes cognitive maps as alternatives to, or additions to, such associations. Their later purpose is planning behavior, not merely storing a record of past actions.
Replanning After Change
A learned representation becomes especially important when the environment or the goal changes unexpectedly. An agent that relies only on a previously learned state-action association may be tied to an earlier response. An agent that can use relationships in an environment model has another possibility: it can use the representation to decide how the changed situation should affect its plan.
Generated example: Suppose an agent has learned that two parts of an environment are connected by more than one route. If one route becomes unavailable, the useful question is not simply which action was rewarded before. The agent can instead use its representation of the environment to consider whether another route still connects the current location with the goal. This illustrates why a model can support flexible planning when circumstances change.
When analyzing a planning behavior, ask two separate questions: what was learned about the environment, and what made that information relevant at the time of action? The source emphasizes that representation learning can occur without reward, while a later goal or unexpected change can make the representation useful for planning.
Common Interpretation Errors
Assuming that no obvious goal-directed behavior means no learning occurred.
Latent learning separates the occurrence of learning from its immediate expression in behavior.
Fix:
Consider whether the animal could later use information from exploration when a goal becomes important.Treating a cognitive map as only a record of past action sequences.
Tolman's interpretation emphasizes relationships represented in the environment model and flexible planning.
Fix:
Ask whether the representation can support a route different from a practiced sequence.Assuming that reward must create the environment representation.
The source identifies the shared property that cognitive-map-like representations can be learned without relying on reward signals.
Fix:
Separate the original learning of the representation from the later goal or reward that makes it relevant.Treating cognitive maps and environment models as identical terms with identical physical mechanisms.
The source presents cognitive map as a psychological term and environment model as a related reinforcement-learning term.
Fix:
Connect them through their shared pattern: learning an environment representation and later using it to plan.
Check Your Understanding
An agent explores an environment without reward or penalty. Later, a goal becomes important, but the route it previously used is no longer available. Explain how latent learning, a cognitive map or environment model, and model-based planning fit together in this situation.
Hints
- Start by separating when the environment representation was learned from when the goal became important.
- Explain why the changed route makes a simple repetition account less adequate.
- State how the learned representation could support consideration of another route.
Answer Structure
What would a strong explanation include?
Latent learning: The agent may have learned during exploration even though its behavior did not immediately show a goal-directed result.
Environment representation: The experience can contribute to a cognitive map or, in reinforcement-learning terms, an environment model representing relationships in the environment.
Changed circumstances: The unavailable route means that repeating a previously reinforced response may not solve the current problem.
Planning: Once the goal is relevant, the agent can use the learned representation to support a different plan, if the representation contains useful information about the changed environment.
The example links hidden learning, representation, motivation, and flexible planning without claiming that psychology and reinforcement learning use identical terminology for identical mechanisms.
Key Takeaways
- Latent learning is learning that is not immediately reflected in behavior.
- Tolman's cognitive map is a learned model of an environment or task space that can support navigation and planning.
- A goal or reward can make an already learned representation relevant without being required for the original learning.
- Cognitive maps and reinforcement-learning environment models share the pattern of representing an environment and later using that representation to plan.
- A learned representation supports flexible behavior, including consideration of a different route when circumstances change.
Key Takeaways
- Latent learning shows that learning and the immediate expression of learning can occur at different times.
- Tolman's cognitive-map perspective explains how exploration can produce a learned environment representation before a goal becomes motivating.
- In reinforcement learning, an environment model provides a related computational way to represent how an environment behaves.
- Model-based planning uses relationships in a learned representation rather than relying only on previously reinforced state-action associations.
- When a route or goal changes unexpectedly, an environment representation can support flexible replanning.