Model-Free and Model-Based Learning
Latent learning occurs without an apparent reward and can be revealed when a reward is later introduced.
Learning Before Reward
Imagine an animal exploring an environment without receiving food or another obvious reward. At first, its behavior may not appear to improve. Later, when a reward becomes available, its performance may improve rapidly. This pattern suggests that learning occurred during the unrewarded period even though no immediate change in behavior was obvious. This kind of learning is called latent learning.
Latent learning is learning that occurs without an apparent reward and can be revealed when a reward is introduced later.
The important distinction is between learning and the visible performance of learned behavior. An animal may acquire useful knowledge before there is a reward that makes the knowledge obvious.
The Maze Experiment
The classic rat-maze experiment is understood in two stages. First, some rats explored a maze without food. During this period, the rats did not receive an apparent reward for navigating it. Second, food was introduced. Once food became available, the rats that had previously explored without food rapidly caught up with rats that had received food throughout the experiment.
Reading the Two Stages
What does rapid improvement after food is introduced suggest about the earlier unrewarded exploration?
Stage 1: The rats explore the maze without food, so there is no apparent reward motivating an immediately visible improvement.
Stage 2: Food becomes available. The previously unrewarded rats then rapidly catch up with rats that had received food throughout the experiment.
Interpretation: The later improvement suggests that useful learning occurred during the earlier exploration and was revealed when the reward appeared.
The absence of an obvious early performance change does not show that no learning occurred.
Beyond Stimulus and Response
A purely stimulus-response account treats learning mainly as the strengthening of direct links between situations and responses. On that view, behavior improves because a particular stimulus comes to trigger a particular response. The maze finding challenged this account because the rats could explore and acquire useful environmental knowledge before food made a particular response immediately rewarding.
The rats' rapid catch-up is consistent with the idea that they had learned something about the maze during exploration. Their later behavior was not adequately described as only repeating a directly rewarded stimulus-response sequence. Instead, earlier experience could be used when the situation changed and a goal became available.
A cognitive map is a mental representation of an environment. The idea is implied by the rats' ability to use earlier maze experience when food later became available.
Two Reinforcement Learning Approaches
Reinforcement learning distinguishes model-free and model-based learning according to whether a model of the environment is used in decision-making. Both approaches concern learning from action outcomes, but they organize what is learned and how it guides later choices differently.
| Approach | Use of an environment model | Decision basis | Behavioral connection |
|---|---|---|---|
| Model-free | Does not use an environment model | Learning from trial and error about what has worked | Habitual behavior |
| Model-based | Uses an environment model | Considering what a choice will lead to | Goal-directed behavior |
Model-free reinforcement learning learns through trial and error without an environment model. It can rely on what has worked before when selecting an action.
Model-based reinforcement learning uses an environment model to make decisions. It can use an understanding of the environment to consider what a choice will lead to.
Habit and Goal Direction
The model-free and model-based distinction aligns with a distinction between habitual and goal-directed behavior. Model-free behavior is associated with relying on what has worked before. With repeated experience, this can support habitual behavior. Model-based behavior is associated with using an understanding of the environment to consider consequences, which supports goal-directed behavior.
Choosing a Route
An agent has learned that one route has produced useful results before. How could the agent use that learning in model-free and model-based ways?
Model-free interpretation: The agent relies on the learned history of what has worked before and selects the associated action without using an environment model.
Model-based interpretation: The agent uses an environment model to consider what different choices will lead to and selects an action in relation to a desired goal.
Behavioral distinction: The first pattern is aligned with habitual behavior. The second is aligned with goal-directed behavior.
The same past experience can support different styles of decision-making depending on whether an environment model is used.
Actor-Critic and Habit Formation
The basic actor-critic method is relevant to habitual behavior because it fits the model-free approach. Model-free algorithms learn through trial and error without an environment model. In an actor-critic description, the actor is associated with action preferences and the critic with value estimates. Repeated reward-based learning can therefore make action selection increasingly depend on learned preferences and value estimates rather than on constructing an explicit model of what each choice will lead to.
The connection to habit is conceptual: repeated trial-and-error learning can make behavior depend on learned action preferences and value estimates. This is why the basic actor-critic method is relevant to habitual behavior rather than to planning with an explicit environment model.
Common Mistakes
Assuming that no visible improvement means no learning occurred.
Latent learning can occur without an apparent reward and may only become visible after a reward is introduced.
Fix:
Separate the occurrence of learning from the immediate performance of learned behavior.Treating the maze finding as ordinary stimulus-response learning only.
The rats that explored without food had not received that apparent reward during the first stage.
Fix:
Recognize that earlier environmental knowledge could be used when food later became available.Defining model-free learning as learning without experience.
Model-free algorithms learn through trial and error.
Fix:
Define model-free learning by the absence of an environment model in decision-making, not by the absence of learning.Calling every learned behavior model-based.
Model-free learning also uses past outcomes, but it does so without an environment model.
Fix:
Ask whether the agent uses an environment model to consider what a choice will lead to.
Check Your Understanding
An animal explores an environment without food. When food is later introduced, its performance improves rapidly. Explain what this demonstrates about latent learning, why the result challenges a purely stimulus-response account, and whether the evidence is more naturally connected to a cognitive map or to a direct stimulus-response sequence.
Hints
- Define latent learning in terms of reward and visible performance.
- Use the two stages of the maze experiment in your explanation.
- Connect a cognitive map to a mental representation of the environment.
Compare these two decision descriptions: an agent chooses an action because that action has worked before; an agent uses an understanding of the environment to consider what each choice will lead to. Identify which description is model-free, which is model-based, and which behavioral tendency each is associated with.
Hints
- Model-free learning does not use an environment model.
- Model-based learning uses an environment model.
- Relate the two approaches to habitual and goal-directed behavior.
Key Takeaways
- Latent learning occurs without an apparent reward and can be revealed when a reward appears later.
- In the classic maze experiment, rats explored without food and later rapidly caught up after food was introduced.
- The finding challenged a purely stimulus-response account and implied that rats could acquire mental representations of the environment, called cognitive maps.
- Model-free reinforcement learning uses trial and error without an environment model and is aligned with habitual behavior.
- Model-based reinforcement learning uses an environment model to consider consequences and is aligned with goal-directed behavior; the basic actor-critic method is relevant to the model-free, habitual side of this distinction.
Key Takeaways
- Latent learning can occur before an apparent reward and may be revealed by rapid later improvement.
- The rat-maze experiment challenged the idea that learning is only a collection of stimulus-response links.
- Cognitive maps are mental representations of environments implied by the rats' use of earlier maze experience.
- Model-free learning uses trial and error without an environment model, whereas model-based learning uses an environment model to make decisions.
- Model-free learning is associated with habitual behavior and basic actor-critic learning; model-based learning is associated with goal-directed behavior.