Reinforcement Learning and Animal Decision Processes
Model-free and model-based learning differ according to whether a model of the environment is used in decision-making.
Two Ways to Use Experience
Suppose an animal has learned that an action tends to produce a useful result. After learning, it can use that experience in two conceptually different ways. It can rely on what has worked before, or it can use an understanding of the environment to consider what a choice will lead to. Reinforcement learning calls these approaches model-free and model-based learning.
Model-free reinforcement learning learns through trial and error without using a model of the environment during decision-making.
Model-based reinforcement learning uses a model of the environment to make decisions.
The Decision-Making Contrast
| Question | Model-free learning | Model-based learning |
|---|---|---|
| Is an environment model used in decision-making? | No | Yes |
| How is learning characterized? | Trial and error | Use of an environment model |
| Behavioral connection | Habitual behavior | Goal-directed behavior |
The distinction is not simply about whether an agent has ever experienced an environment. Both approaches can learn from experience. The key question is what the agent uses when deciding: a trial-and-error history without an environment model, or an environment model that supports consideration of what an action will lead to.
Tracing an Environment Model
In model-based decision-making, the environment model is the defining extra ingredient. The agent does not merely repeat an action because that action worked before. It uses the model to consider what a choice will lead to and then uses those predicted consequences to support a decision.
Choosing Between Two Actions
An animal is in a situation where action A has produced a useful result in the past. Consider how the decision can be described under the two reinforcement-learning approaches.
Model-free description: The animal can use trial-and-error learning without an environment model and rely on the action that has worked before.
Model-based description: The animal can use an environment model to consider what action A or another action will lead to before choosing.
Identify the distinction: The important difference is not the action's history alone. It is whether the decision uses an environment model.
Model-free learning relies on learned experience without an environment model, whereas model-based learning uses an environment model to support the choice.
From Repeated Success to Habit
Model-free learning is related to habitual behavior because the learned experience can guide action without an environment model being used to consider future consequences. In the source description, an animal may rely on what has worked before. This is the habitual side of the model-free and model-based distinction.
When Goals Change the Choice
Model-based learning is related to goal-directed behavior because the agent can use an environment model to consider what a choice will lead to. If the useful result or goal changes, the model-based description allows the agent to reconsider the consequences of available actions rather than relying only on which action worked previously.
Imagine that an animal has two possible actions and that the consequences of those actions differ. A model-free description emphasizes the action that has tended to produce a useful result. A model-based description emphasizes using an environment model to consider the consequences in relation to the current goal. The example illustrates the conceptual connection between model-based learning and goal-directed behavior; it does not require the agent simply to repeat its past choice.
The Actor-Critic Connection
The basic actor-critic method is relevant to habitual behavior because the model-free side of reinforcement learning is based on trial-and-error learning without an environment model. When actor-critic learning is discussed in this context, its importance is the connection between learning from experienced results and behavior that can rely on what has worked before, rather than requiring an explicit model of the environment for each decision.
Common Classification Errors
Treating model-free learning as learning without experience
Model-free algorithms learn through trial and error. Their defining feature is the absence of an environment model in decision-making, not the absence of learning.
Fix:
Describe model-free learning as experience-based trial-and-error learning without an environment model.Defining model-based learning as simply remembering past rewards
The defining feature of model-based learning is using an environment model to make decisions.
Fix:
Ask whether the agent uses a model of the environment to consider what a choice will lead to.Assuming habitual and goal-directed are unrelated to reinforcement learning
The source explicitly aligns model-free learning with habitual behavior and model-based learning with goal-directed behavior.
Fix:
Use the behavioral terms as conceptual connections: model-free with habitual behavior, and model-based with goal-directed behavior.Thinking a model-based decision must repeat the action that worked before
Model-based learning uses an environment model to consider what choices will lead to.
Fix:
Focus on predicted consequences and the current goal rather than on repetition alone.
Check Your Understanding
An animal repeatedly chooses an action because that action has tended to produce a useful result. No model of the environment is used while choosing. Classify this description and explain its connection to behavior.
Hints
- Look for whether an environment model is used during decision-making.
- Then connect the classification to habitual or goal-directed behavior.
What do you think happens?
An agent uses an environment model to consider what different actions will lead to. Is this model-free or model-based learning?
Reveal answer
Answer: Model-based
The defining distinction is the use of an environment model in decision-making.
Classifying a Decision Process
A learner describes an animal as using its past experience to select an action, without consulting a model of what the environment will do next.
Find the decision rule: The description says that the action is selected from past experience.
Check for a model: The description explicitly says that no environment model is consulted.
Connect the behavior: This matches model-free learning and its conceptual connection with habitual behavior.
The process is model-free reinforcement learning, related to habitual behavior.
Key Takeaways
- Model-free and model-based reinforcement learning differ according to whether a model of the environment is used in decision-making.
- Model-free algorithms learn through trial and error without an environment model.
- Model-based algorithms use an environment model to consider what choices will lead to.
- The distinction aligns model-free learning with habitual behavior and model-based learning with goal-directed behavior.
- The basic actor-critic method is relevant to habitual behavior because it is connected here with trial-and-error learning that does not require an environment model for each decision.