Concepts / Dopamine as a Reinforcement Signal

Dopamine as a Reinforcement Signal

The actor-critic model is mapped hypothetically onto different striatal subdivisions rather than onto one undivided structure.

  • Programming

From Actor-Critic Model to Brain Circuit

The actor-critic model separates two jobs. The actor influences which action is selected, while the critic learns value information. A proposed neural mapping assigns these jobs to different subdivisions of the striatum rather than treating the striatum as one undivided structure. This is a hypothesis and a schematic analogy, not a claim that the artificial model and the brain are identical.

proposed correspondencevalue-learning partActorAction selectionDorsal striatumInfluences action selectionCriticValue learningVentral striatumValue and reward processing
Which striatal subdivisions are associated with action selection and value learning in the proposed mapping?

The division is a division of emphasis, not a claim that one subdivision receives only one kind of information. The dorsal striatum is primarily associated with influencing action selection. The ventral striatum is associated with value learning and reward processing, including assigning affective value to sensations. In this proposed mapping, the actor is associated with the dorsal striatum, while the value-learning part of the critic is associated with the ventral striatum.

Value and Reward Become Dopamine Activity

To reconstruct the proposed mechanism, begin with information about a situation. The cerebral cortex, together with other structures, sends the striatum information about stimuli, internal states, and motor activity. The ventral striatum supplies value information to dopamine neurons in the ventral tegmental area and substantia nigra pars compacta. These dopamine neurons combine value information with information about reward. Their resulting activity corresponds to the temporal-difference error signal in the model.

information reachesvalue informationreward informationactivity corresponds toSituationinformationStimuli, internal states,motor activityVentral striatumValue informationDopamine neuronsVTA and SNpcTD error signalProposed neuralinterpretationReward informationReward-related input
How do situation information, value, and reward contribute to the proposed dopamine signal?

Tracing One Proposed Signal

Trace the proposed route from information about a situation to reinforcement-related dopamine activity.

Situation information: Information about stimuli, internal states, and motor activity reaches the striatum from the cerebral cortex and other structures.

Value processing: The ventral striatum is associated with value learning and reward processing, so it supplies value information in the proposed circuit.

Dopamine-neuron input: Value information from the ventral striatum reaches dopamine neurons in the VTA and SNpc, along with reward-related information.

Reinforcement-related activity: The resulting dopamine-neuron activity corresponds to the model's temporal-difference error signal and can provide reinforcement for learning.

The proposal links situation information, ventral-striatal value information, reward information, and dopamine activity in one circuit-level explanation.

Dopamine at Cortical-Striatal Synapses

Dopamine is not presented as an isolated output. Dopamine neurons have widely branching axons, and those axons make synaptic contact with spines on the dendrites of medium spiny neurons. Medium spiny neurons are the main input/output neurons of both the dorsal and ventral striatum. Cortical axons also contact the tips of these spines.

cortical contactdopamine contactpart ofsupportsCortical axonInput to a spineDendritic spineContact siteDopamine axonBranching inputMedium spiny neuronDorsal or ventral striatumSynaptic learningEffectiveness can change
Where does dopamine act relative to cortical input and striatal neurons, and why can timing matter for learning?

The proposed learning rule is located at the relationship between cortical input and striatal neurons. Cortical axons provide input at dendritic spines, while dopamine axons also contact those spines. According to the hypothesis, the effectiveness of cortical-to-striatal synapses can change in a way that depends critically on dopamine reinforcement. The timing and presence of dopamine therefore provide a way for an outcome-related signal to influence learning at connections involved in striatal processing.

From Dopamine Activity to Repeated Action

A reinforcement signal helps explain why an animal repeats one behavior instead of another. In the proposed account, the dorsal striatum is associated with influencing action selection, while dopamine activity supplies reinforcement that can affect learning at striatal synapses. If dopamine activity follows an instrumental behavior, the learning process can make the relationship between the relevant cortical input and striatal response more effective. The behavior can then become more likely to be selected again.

supports selectionfollowed byreinforces learning atinfluences laterSituationCortical and otherinformationInstrumental actionSelected behaviorPhasic dopamineactivityReinforcement signalCortical-striatalsynapseLearning can changeeffectivenessAction selectionAction may be repeated
How can phasic dopamine activity following an action contribute to repeating that action?

Evidence from Self-Stimulation

In the Olds and Milner experiment from 1954, rats could press a lever to receive electrical stimulation. The rats learned to press the lever, demonstrating that electrical stimulation of relevant brain systems could reinforce an instrumental behavior. This experiment supplied an early behavioral demonstration that an animal may repeat an action when that action produces a reinforcing brain event.

responseproducesreinforcesloop continuesLever availableRat can respondLever pressInstrumental behaviorElectricalstimulationBrain-system stimulationRepeated pressingLearned response
What behavioral pattern emerged when rats could press a lever to produce electrical stimulation?

The important inference is behavioral: the rats learned an instrumental response because the response produced electrical stimulation. The experiment did not by itself establish every detail of the dopamine actor-critic mapping. It showed that stimulation of brain systems can function as a reinforcing consequence for lever pressing.

Optogenetic Evidence and Species Limits

Optogenetic experiments provide causal evidence by using light-sensitive proteins and laser flashes to control selected dopamine neurons. These experiments show that phasic dopamine activation can produce learned preferences and influence whether responding and learning continue. This supports the interpretation that phasic dopamine activity can act like a reinforcement signal in control and prediction.

included in evidence baseincluded in evidence baseincluded in evidence baseMicePhasic dopamine activationstudiedReinforcementevidenceSpecies-specific outcomesnot detailedRatsPhasic dopamine activationstudiedFruit fliesPhasic dopamine activationstudied
What can be distinguished about phasic dopamine activation in mice, rats, and fruit flies from the provided material?
  • Treating the actor-critic mapping as a proven one-to-one identity between an artificial model and the brain.

    The mapping is explicitly proposed and schematic. It assigns different emphases to striatal subdivisions rather than claiming that the brain and model are identical.

    Fix: Describe the dorsal striatum as associated with action selection and the ventral striatum as associated with value learning and reward processing in the proposed mapping.

  • Saying that dopamine activity is proven to be exactly the temporal-difference error.

    The source says dopamine activity corresponds to the model's signal and acts like a reinforcement signal, but does not claim exact identity.

    Fix: Use the more careful wording that dopamine activity is a proposed neural counterpart or correspondence to the temporal-difference error signal.

  • Claiming that the Olds and Milner experiment directly proved the complete dopamine mechanism.

    The experiment demonstrated reinforcement by electrical stimulation and does not, by itself, establish every part of the proposed circuit.

    Fix: Use it as behavioral evidence that brain stimulation can reinforce an instrumental response.

  • Inventing different mouse, rat, and fruit-fly outcomes when the supplied material does not list them.

    The source pack mentions the species but does not provide enough detail for a species-by-species comparison.

    Fix: State only that the evidence base includes those species and that the supplied material does not specify the differences.

Check Your Understanding

MEDIUM

Explain the proposed route from a situation to repeated action. Your explanation should name the dorsal striatum, ventral striatum, dopamine neurons in the VTA and SNpc, the temporal-difference error correspondence, and cortical-to-striatal synapses.

Hints
  • Start with the different emphases assigned to the dorsal and ventral striatum.
  • Then explain what information reaches dopamine neurons.
  • Finish by describing how dopamine can influence learning at striatal spines and later action selection.

What do you think happens?

A rat can press a lever, and each press produces electrical stimulation of a relevant brain system. What behavior would count as evidence that the stimulation was reinforcing?

  • The rat learns to press the lever repeatedly
  • The rat never presses the lever
  • The lever becomes physically heavier
  • The rat loses the ability to receive stimulation
Reveal answer

Answer: The rat learns to press the lever repeatedly.

The Olds and Milner experiment demonstrated that rats learned lever pressing when the response produced electrical stimulation, showing that the stimulation could reinforce an instrumental behavior.

Key Takeaways

  • The proposed actor-critic mapping associates the dorsal striatum with action selection and the ventral striatum with value learning and reward processing.
  • The ventral striatum is proposed to send value information to dopamine neurons in the VTA and SNpc, where value and reward information contribute to activity corresponding to a temporal-difference error signal.
  • Dopamine axons contact dendritic spines of medium spiny neurons, where dopamine may influence learning at cortical-to-striatal synapses.
  • The Olds and Milner experiment showed that rats learned lever pressing when pressing produced electrical stimulation, supporting a reinforcing role for brain stimulation.
  • Optogenetic studies support a reinforcing role for phasic dopamine activity, but the supplied material does not specify distinct effects for mice, rats, and fruit flies.