Concepts / Model-Based Reinforcement Learning in the Brain

Model-Based Reinforcement Learning in the Brain

Goal-directed and habitual behavior rely on different, though not necessarily exclusive, brain contributions.

  • Programming

A Choice That Changes

Imagine an animal that has learned an action leading to a valuable outcome. If that outcome becomes less valuable, a flexible animal should reconsider the action. This adjustment is characteristic of goal-directed behavior. Habitual behavior is different: an established action can continue even when its outcome is no longer as valuable. The evidence suggests that these two modes of control rely on different, though not necessarily exclusive, brain contributions.

What do you think happens?

An animal keeps performing a learned action after the outcome becomes less valuable. Which control mode does this behavior resemble more closely?

  • Goal-directed control
  • Habitual control
Reveal answer

Answer: Habitual control

Goal-directed behavior should be sensitive to changes in outcome value, whereas an established habit can continue even when the outcome is no longer as valuable.

Two Ways to Select an Action

Model-based and habitual control represent different kinds of guidance for behavior. In model-based behavior, the animal uses information about the environment and expected outcomes to support a flexible choice. In habitual behavior, a learned action can be expressed without the same degree of adjustment when outcome value changes. The contrast is functional rather than a claim that every decision follows a simple one-way sequence.

evaluatechoose using expected valueretrieve learned responserepeat actionCurrent situationExpected outcomevalue can changeGoal-directed choicemodel-basedEstablished actionhabitHabitual choicemodel-free
How does the brain choose between using expected outcomes and relying on an established action?

The model-based side is linked more strongly with goal-directed behavior and the dorsomedial striatum. The model-free side is linked more strongly with habit learning and the dorsolateral striatum.

Dorsal Striatum Division

The dorsal striatum includes the dorsolateral striatum, or DLS, and the dorsomedial striatum, or DMS. The DLS is more closely associated with model-free learning and habits. The DMS is more closely associated with model-based learning and goal-directed processes.

associated withassociated withDorsolateralstriatumDLSHabit learningmodel-freeDorsomedialstriatumDMSGoal-directedlearningmodel-based
What is the functional contrast between the two striatal regions?

Reading an Inactivation Result

What would a temporary inactivation experiment suggest if inactivating one striatal region weakens habit learning and the animal becomes more goal directed?

Identify the changed behavior: Habit learning is impaired, while behavior becomes more goal directed.

Connect behavior to the region: The region is implicated more strongly in habitual control.

Apply the striatal contrast: This pattern is associated with the DLS, which is more closely linked to habits and model-free learning.

The result supports a stronger role for the DLS in habitual control.

Value, Planning, and Memory

Goal-directed choice depends on more than one brain contribution. The prefrontal cortex is implicated in executive functions such as planning and decision making. Within the prefrontal cortex, the orbitofrontal cortex, or OFC, is related both to the subjective reward value of biologically significant stimuli and to the reward expected as a consequence of actions. This makes the OFC a candidate contributor to the reward portion of an animal's environment model.

consider expected consequenceprovide reward informationguide selectionPossible actionOrbitofrontal cortexsubjective and expectedrewardPlanning and decisionmakingprefrontal contributionGoal-directed choice
How can prefrontal contributions provide information used in selecting an action?

The hippocampus adds a different kind of support. It is critical for memory and spatial navigation. In rats, hippocampal function is important for navigating a maze in a goal-directed manner. This connects the hippocampus with the possibility that animals use models, or cognitive maps, when selecting actions. The hippocampus may also contribute to the human ability to imagine new experiences.

supportssupportscontributes tocontributes tosupportsHippocampusMemoryCognitive mapenvironment modelModel-based planningSpatial navigation
How are memory and spatial representations connected to planning and model-based behavior?

Changing the Balance

Temporary inactivation experiments provide a way to test how changing activity in a region affects behavioral control. When the DLS is inactivated, habit learning is impaired and the animal relies more on goal-directed processes. When the DMS is inactivated, goal-directed processes are impaired and the animal relies more on habit learning.

reduce DLS activitybehavior shifts towardreduce DMS activitybehavior shifts towardBehavioral controlhabitual and goal-directedcontributionsDLS inactivatedMore goal-directedbehaviorhabit learning impairedDMS inactivatedMore habitualbehaviorgoal-directed processesimpaired
What happens to behavioral control when activity in a region supporting habits or goal-directed processes is reduced?

Predicting the Direction of a Shift

Predict the behavioral effect of temporarily inactivating the DMS.

Recall the DMS association: The DMS is more closely associated with model-based learning and goal-directed processes.

Identify the impaired process: Inactivating the DMS impairs goal-directed processes.

Predict the compensating pattern: The animal relies more on habit learning, so behavior shifts toward habitual control.

DMS inactivation is expected to produce more habitual behavior and weaker goal-directed behavior.

Testing the Mechanism

A common way to investigate these roles is to temporarily inactivate a brain area in a rat and then observe behavior in an outcome-devaluation experiment. The logic is comparative. If inactivating a region weakens habit learning while behavior becomes more goal directed, that region is implicated more strongly in habitual control. If inactivating another region weakens goal-directed processes while behavior becomes more habitual, that region is implicated more strongly in goal-directed control.

ManipulationObserved behavioral patternInterpretation
DLS inactivationHabit learning is impaired; behavior becomes more goal directedThe DLS is more strongly implicated in habitual control
DMS inactivationGoal-directed processes are impaired; behavior becomes more habitualThe DMS is more strongly implicated in goal-directed control

Comparative interpretation of temporary striatal inactivation

  • Treating the DLS and DMS contrast as an absolute separation

    The evidence describes different, though not necessarily exclusive, contributions and presents the contrast as a way to organize the findings.

    Fix: Use qualified language: the DLS is more closely associated with habits and model-free learning, while the DMS is more closely associated with goal-directed and model-based processes.

  • Reducing the OFC to a general movement controller

    The source emphasizes the OFC's relationship to subjective reward value and reward expected as a consequence of actions.

    Fix: Explain the OFC as a contributor of reward-value information relevant to an environment model and goal-directed choice.

  • Treating the hippocampus as relevant only to memory

    The hippocampus contributes to both memory and spatial navigation, and hippocampal function is important for goal-directed maze navigation in rats.

    Fix: Connect memory and spatial navigation to the use of models or cognitive maps during action selection.

Apply the Contrast

MEDIUM

An experiment temporarily inactivates a brain region. Afterward, an animal is less able to adjust its behavior when the value of an expected outcome changes and instead continues an established action. Which behavioral mode has become more prominent, and which striatal region is more closely associated with that mode?

Hints
  • Ask whether the animal is responding flexibly to outcome value or repeating a learned action.
  • Use the DLS and DMS contrast to identify the associated region.
MEDIUM

An animal must navigate a maze toward a goal and use information about expected reward when selecting an action. Name two brain contributions that could support this model-based choice, and state what kind of information each contributes.

Hints
  • One contribution is associated with reward value and expected reward.
  • Another contribution is associated with memory and spatial navigation.

A strong answer should identify habitual control and the DLS for the first prompt. For the second prompt, it should identify the OFC as a source of subjective and expected reward information and the hippocampus as a contributor to memory and spatial navigation.

Working Model

  1. The DLS is more closely associated with model-free learning and habitual behavior.
  2. The DMS is more closely associated with model-based learning and goal-directed behavior.
  3. The prefrontal cortex contributes to planning and decision making, while the OFC provides information about subjective and expected reward.
  4. The hippocampus supports memory and spatial navigation, linking it to cognitive maps and model-based behavior.
  5. Inactivation results help reveal functional contributions: reducing DLS activity shifts behavior toward goal-directed processes, while reducing DMS activity shifts behavior toward habits.

Key Takeaways

  • Goal-directed behavior can adjust when the value of an expected outcome changes, whereas habitual behavior can continue despite reduced outcome value.
  • The DLS is more closely associated with habitual, model-free control, and the DMS is more closely associated with goal-directed, model-based control.
  • The OFC contributes subjective and expected reward information, while the broader prefrontal cortex supports planning and decision making.
  • The hippocampus contributes memory and spatial navigation that can support cognitive maps and model-based behavior.
  • Temporary inactivation experiments predict opposite behavioral shifts: DLS inactivation favors more goal-directed behavior, while DMS inactivation favors more habitual behavior.