Habit Learning and Goal-Directed Behavior
Goal-directed and habitual behavior rely on different, though not necessarily exclusive, brain contributions.
A Choice That Stops Updating
Imagine an animal that has learned an action usually produces a valuable outcome. Later, the outcome becomes less valuable. A flexible, goal-directed animal should reconsider the action because its current goal has changed. A habitual animal may continue performing the established action. This contrast gives researchers a way to study how behavior is controlled.
Habitual and goal-directed behavior are different but not necessarily exclusive modes of control. Evidence points to different contributions from parts of the dorsal striatum, prefrontal cortex, and hippocampus.
Outcome Value as a Behavioral Test
A devalued outcome
An animal has learned that Action A leads to Outcome X. Outcome X later becomes less valuable. What should happen if the action is controlled mainly by a current goal, and what might happen if it is controlled mainly by a habit?
Initial learning: The animal learns a relationship between Action A and Outcome X.
Change in value: The value of Outcome X changes, creating a reason to reconsider Action A.
Goal-directed prediction: If the animal uses current outcome value, it should reduce or reconsider Action A.
Habitual prediction: If the action is controlled by an established habit, the animal may continue Action A even though Outcome X is less valuable.
Sensitivity to the changed outcome supports goal-directed control; persistence despite the changed outcome supports habitual control.
Outcome-devaluation experiments are diagnostic rather than merely tests of whether an animal can perform an action. They ask what information appears to control the action: current information about the outcome's value, or information about choosing the action that has already been stored.
Two Parts of the Dorsal Striatum
| Region | Closer association | Behavioral implication of temporary inactivation |
|---|---|---|
| Dorsolateral striatum, or DLS | Model-free learning and habits | Habit learning is impaired, and behavior relies more on goal-directed processes |
| Dorsomedial striatum, or DMS | Model-based learning and goal-directed processes | Goal-directed processes are impaired, and behavior relies more on habit learning |
The dorsal striatum is divided into the DLS and DMS. Temporary inactivation experiments reveal a functional contrast: reducing DLS activity impairs habit learning and leaves greater reliance on goal-directed processes, whereas reducing DMS activity impairs goal-directed processes and leaves greater reliance on habit learning. These findings support a division of contribution, not a claim that every decision follows a simple one-way pathway.
Cached Values and Environmental Models
Model-free and model-based describe how an action is selected. Model-free control uses cached policy or action-value information: information about choosing actions has already been stored through learning. This is associated more strongly with habitual behavior and the DLS. Model-based control plans ahead with an environmental model. It can use information about the environment and expected outcomes to select an action, so it is associated more strongly with goal-directed behavior and the DMS.
Prefrontal Reward Information
The prefrontal cortex is the front-most part of the frontal cortex and is implicated in executive functions such as planning and decision making. Within it, the orbitofrontal cortex lies immediately above the eyes. OFC activity is related both to the subjective reward value of biologically significant stimuli and to the reward expected as a consequence of actions.
Because the OFC carries information related to subjective and expected reward, it is a candidate contributor to the reward portion of an animal's environmental model. That information can support goal-directed choice by helping relate an action to the reward expected from its consequence. The OFC is therefore part of the broader network contributing information to a choice, rather than a complete explanation of all goal-directed behavior.
Memory, Maps, and Navigation
The hippocampus contributes memory and spatial-navigation abilities that support model-based behavior. In rats, hippocampal function is important for navigating a maze in a goal-directed manner. This connects the hippocampus with the idea that animals can use models, or cognitive maps, when selecting actions. The hippocampus may also contribute to the human ability to imagine new experiences.
The hippocampus contributes a different kind of support from the OFC. The OFC is associated with reward value and expected reward, whereas the hippocampus contributes memory and spatial-navigation information that can help an animal use an environmental model.
Predicting a Regional Shift
What do you think happens?
An experiment temporarily inactivates the DLS. Which behavioral shift is most consistent with the evidence?
Reveal answer
Answer: Habit learning is impaired and behavior relies more on goal-directed processes.
The DLS is more closely associated with model-free learning and habits. The source evidence states that temporary DLS inactivation impairs habit learning and produces greater reliance on goal-directed processes.
What do you think happens?
An experiment temporarily inactivates the DMS. Which behavioral shift is most consistent with the evidence?
Reveal answer
Answer: Goal-directed processes are impaired and behavior relies more on habit learning.
The DMS is more closely associated with model-based learning and goal-directed processes. Temporary DMS inactivation impairs goal-directed processes and increases reliance on habit learning.
A behavior continues after its outcome becomes less valuable. Explain which type of control this pattern is more consistent with, and identify the kind of information that may be guiding the action.
Hints
- Ask whether the action is responding to the current value of the outcome.
- Compare cached action-value information with planning using an environmental model.
Treating habits and goal-directed behavior as completely separate systems that never overlap.
The source describes different, though not necessarily exclusive, brain contributions and presents the distinction as a functional contrast.
Fix:
Use language such as more closely associated with or greater reliance on a control process.Defining model-free learning as learning without any stored information.
Model-free control is associated with cached policy or action-value information.
Fix:
Describe it as selecting actions using previously stored decision information rather than planning ahead with an environmental model.Assuming that outcome devaluation directly proves the identity of one single brain region.
Outcome-devaluation experiments provide evidence about the type of control guiding behavior, while inactivation experiments help compare regional contributions.
Fix:
Treat the result as diagnostic evidence about habitual or goal-directed control and interpret it alongside regional manipulations.Assigning all model-based behavior to the DMS alone.
The source describes several systems contributing different information to a goal-directed choice.
Fix:
Relate the DMS to model-based processes while also considering OFC reward information and hippocampal memory and spatial-navigation support.
Putting the Control Modes Together
- The DLS is more closely associated with model-free learning and habitual behavior; temporary DLS inactivation impairs habit learning and shifts reliance toward goal-directed processes.
- The DMS is more closely associated with model-based learning and goal-directed behavior; temporary DMS inactivation impairs goal-directed processes and shifts reliance toward habits.
- The OFC contributes information about subjective reward value and expected reward, making it relevant to the reward component of an environmental model.
- The hippocampus contributes memory and spatial-navigation abilities that support model-based behavior, including goal-directed maze navigation.
- Outcome-devaluation experiments help researchers determine whether behavior is sensitive to current outcome value or continues through cached habitual control.
Key Takeaways
- Habitual behavior is more closely associated with cached, model-free action information and the DLS.
- Goal-directed behavior is more closely associated with model-based planning and the DMS.
- The OFC contributes subjective and expected reward information, while the hippocampus contributes memory and spatial-navigation support.
- Outcome devaluation tests whether behavior updates when the value of an outcome changes.