Model-Based Learning and Decision Making
Goal-directed behavior can use model-based simulation of future actions, while habit behavior is associated with model-free learning.
Two Routes to Action
A behavior can be guided in at least two fundamentally different ways. Goal-directed behavior can use a model of the environment to simulate possible future courses of action before choosing. Habit behavior is associated with model-free learning, in which behavior is guided by information learned from previous reward experience rather than by explicitly simulating future possibilities.
Simulating a Possible Future
In model-based, goal-directed control, the learner uses a model of the environment to consider what could happen after different actions. The decision process therefore includes a future-simulation step: possible actions are considered, their possible consequences are simulated, and an action is selected after those consequences have been taken into account.
Choosing after imagining consequences
A learner has two possible actions and uses an environment model to consider what could happen after each one.
Represent the environment: The learner uses a model of the environment rather than relying only on a previously learned action tendency.
Simulate the alternatives: The learner considers possible future courses of action and the outcomes that could follow them.
Select an action: The learner chooses after the simulated consequences have been considered.
This is a model-based description of goal-directed control: the future simulation is part of the route to the decision.
Tracing the Decision State
The important state change in a model-based decision is the movement from a set of possible actions to simulated future consequences and then to a selected action. The learner is not described as merely repeating an action because it was rewarded before. Instead, the learner uses information about how the environment works to consider possible futures.
Learning Without Explicit Simulation
Habit behavior is associated with model-free learning. In this description, behavior is guided by information learned from previous reward experience rather than by explicitly simulating future possibilities at the time of the decision. The key contrast is therefore not that habits contain no learned information. They do contain information from past reward experience, but the action is not selected through an explicit simulation of possible futures.
Comparing Control Strategies
| Feature | Model-based, goal-directed control | Model-free, habitual control |
|---|---|---|
| Main source of guidance | A model of the environment | Information learned from previous reward experience |
| Role of future consequences | Possible futures are explicitly simulated | Action is not selected through explicit future simulation |
| Behavioral description | Goal-directed | Habit-associated |
| What the contrast does not establish | It does not identify one completely separate neural structure | It does not prove that every reward-related signal is purely model-free |
The Unresolved Arbitration Problem
The difficult question is not simply whether both forms of control exist. The difficult question is how the brain decides which kind of control to use. A decision may involve goal-directed and habitual influences, but the brain's method for arbitrating between them remains unknown.
This is an open neural question because the computational difference is sharper than the neural evidence. The source does not provide a positive answer that model-based and model-free systems are cleanly separated in the brain. Arbitration is therefore a future-direction topic: a complete account must explain both the computational difference and how their influences are combined or selected.
Distributed Neural Influences
Model-based influences cannot be assumed to occur only in regions identified with model-based learning. The evidence summarized in the source indicates that model-based influences can appear broadly wherever the brain processes reward information, including regions thought to be important for model-free learning.
Common Interpretation Mistakes
Treating model-based and model-free learning as two completely separate brain locations.
The source states that computationally distinct systems do not necessarily correspond to cleanly separated neural structures.
Fix:
Keep the computational distinction, but allow for overlapping or distributed neural influences.Assuming that a reward-related signal is automatically model-free.
The source notes that dopamine signals can also show the influence of model-based information.
Fix:
Treat the signal as evidence requiring interpretation rather than as proof of a purely model-free process.Defining model-free learning as learning with no useful information.
Model-free behavior is guided by information learned from previous reward experience.
Fix:
Describe model-free learning as guidance from past reward information without explicit future simulation at the time of choice.Explaining the existence of both systems without addressing arbitration.
The arbitration mechanism is the unresolved neural question highlighted by the source.
Fix:
Separate the question of whether both forms of control exist from the question of how control is selected or combined.
Check Your Understanding
A learner chooses an action after using a model of the environment to consider possible future consequences. Identify the control strategy and explain why it fits that strategy.
Hints
- Look for whether future possibilities are explicitly simulated.
- Connect the use of an environment model with goal-directed behavior.
A researcher observes a reward-related dopamine signal in a region commonly associated with model-free learning. What conclusion is justified, and what conclusion is not justified?
Hints
- The signal may be related to reward processing.
- Do not assume that its location proves the signal is purely model-free.
Explain why the question of arbitration remains open even after distinguishing model-based and model-free learning computationally.
Hints
- The computational distinction does not guarantee cleanly separated neural structures.
- Include the problem of deciding which form of control to use.
Key Takeaways
- Model-based goal-directed behavior can use an environment model to simulate possible future actions and outcomes before choosing.
- Model-free learning is associated with habit behavior guided by information learned from previous reward experience rather than by explicit future simulation.
- The brain's method for arbitrating between goal-directed and habitual control remains unknown.
- Computationally distinct learning systems do not necessarily map onto cleanly separated neural structures.
- Model-based influences can appear in reward-processing regions and dopamine signals commonly associated with model-free learning, so neural location alone does not establish a purely model-free process.
Key Takeaways
- Goal-directed control can simulate future consequences using a model of the environment.
- Habit behavior is associated with model-free learning from previous reward experience.
- Arbitration between goal-directed and habitual control is an unresolved neural question.
- Model-based influences may appear in reward-processing regions and dopamine signals associated with model-free learning.
- The computational distinction between learning systems should not be mistaken for a clean anatomical separation.