Dopamine as a Reinforcement Signal
The actor-critic model is mapped hypothetically onto different striatal subdivisions rather than onto one undivided structure.
From Actor-Critic Model to Brain Circuit
The actor-critic model separates two jobs. The actor influences which action is selected, while the critic learns value information. A proposed neural mapping assigns these jobs to different subdivisions of the striatum rather than treating the striatum as one undivided structure. This is a hypothesis and a schematic analogy, not a claim that the artificial model and the brain are identical.
The division is a division of emphasis, not a claim that one subdivision receives only one kind of information. The dorsal striatum is primarily associated with influencing action selection. The ventral striatum is associated with value learning and reward processing, including assigning affective value to sensations. In this proposed mapping, the actor is associated with the dorsal striatum, while the value-learning part of the critic is associated with the ventral striatum.
Value and Reward Become Dopamine Activity
To reconstruct the proposed mechanism, begin with information about a situation. The cerebral cortex, together with other structures, sends the striatum information about stimuli, internal states, and motor activity. The ventral striatum supplies value information to dopamine neurons in the ventral tegmental area and substantia nigra pars compacta. These dopamine neurons combine value information with information about reward. Their resulting activity corresponds to the temporal-difference error signal in the model.
Tracing One Proposed Signal
Trace the proposed route from information about a situation to reinforcement-related dopamine activity.
Situation information: Information about stimuli, internal states, and motor activity reaches the striatum from the cerebral cortex and other structures.
Value processing: The ventral striatum is associated with value learning and reward processing, so it supplies value information in the proposed circuit.
Dopamine-neuron input: Value information from the ventral striatum reaches dopamine neurons in the VTA and SNpc, along with reward-related information.
Reinforcement-related activity: The resulting dopamine-neuron activity corresponds to the model's temporal-difference error signal and can provide reinforcement for learning.
The proposal links situation information, ventral-striatal value information, reward information, and dopamine activity in one circuit-level explanation.
Dopamine at Cortical-Striatal Synapses
Dopamine is not presented as an isolated output. Dopamine neurons have widely branching axons, and those axons make synaptic contact with spines on the dendrites of medium spiny neurons. Medium spiny neurons are the main input/output neurons of both the dorsal and ventral striatum. Cortical axons also contact the tips of these spines.
The proposed learning rule is located at the relationship between cortical input and striatal neurons. Cortical axons provide input at dendritic spines, while dopamine axons also contact those spines. According to the hypothesis, the effectiveness of cortical-to-striatal synapses can change in a way that depends critically on dopamine reinforcement. The timing and presence of dopamine therefore provide a way for an outcome-related signal to influence learning at connections involved in striatal processing.
From Dopamine Activity to Repeated Action
A reinforcement signal helps explain why an animal repeats one behavior instead of another. In the proposed account, the dorsal striatum is associated with influencing action selection, while dopamine activity supplies reinforcement that can affect learning at striatal synapses. If dopamine activity follows an instrumental behavior, the learning process can make the relationship between the relevant cortical input and striatal response more effective. The behavior can then become more likely to be selected again.
Evidence from Self-Stimulation
In the Olds and Milner experiment from 1954, rats could press a lever to receive electrical stimulation. The rats learned to press the lever, demonstrating that electrical stimulation of relevant brain systems could reinforce an instrumental behavior. This experiment supplied an early behavioral demonstration that an animal may repeat an action when that action produces a reinforcing brain event.
The important inference is behavioral: the rats learned an instrumental response because the response produced electrical stimulation. The experiment did not by itself establish every detail of the dopamine actor-critic mapping. It showed that stimulation of brain systems can function as a reinforcing consequence for lever pressing.
Optogenetic Evidence and Species Limits
Optogenetic experiments provide causal evidence by using light-sensitive proteins and laser flashes to control selected dopamine neurons. These experiments show that phasic dopamine activation can produce learned preferences and influence whether responding and learning continue. This supports the interpretation that phasic dopamine activity can act like a reinforcement signal in control and prediction.
Treating the actor-critic mapping as a proven one-to-one identity between an artificial model and the brain.
The mapping is explicitly proposed and schematic. It assigns different emphases to striatal subdivisions rather than claiming that the brain and model are identical.
Fix:
Describe the dorsal striatum as associated with action selection and the ventral striatum as associated with value learning and reward processing in the proposed mapping.Saying that dopamine activity is proven to be exactly the temporal-difference error.
The source says dopamine activity corresponds to the model's signal and acts like a reinforcement signal, but does not claim exact identity.
Fix:
Use the more careful wording that dopamine activity is a proposed neural counterpart or correspondence to the temporal-difference error signal.Claiming that the Olds and Milner experiment directly proved the complete dopamine mechanism.
The experiment demonstrated reinforcement by electrical stimulation and does not, by itself, establish every part of the proposed circuit.
Fix:
Use it as behavioral evidence that brain stimulation can reinforce an instrumental response.Inventing different mouse, rat, and fruit-fly outcomes when the supplied material does not list them.
The source pack mentions the species but does not provide enough detail for a species-by-species comparison.
Fix:
State only that the evidence base includes those species and that the supplied material does not specify the differences.
Check Your Understanding
Explain the proposed route from a situation to repeated action. Your explanation should name the dorsal striatum, ventral striatum, dopamine neurons in the VTA and SNpc, the temporal-difference error correspondence, and cortical-to-striatal synapses.
Hints
- Start with the different emphases assigned to the dorsal and ventral striatum.
- Then explain what information reaches dopamine neurons.
- Finish by describing how dopamine can influence learning at striatal spines and later action selection.
What do you think happens?
A rat can press a lever, and each press produces electrical stimulation of a relevant brain system. What behavior would count as evidence that the stimulation was reinforcing?
Reveal answer
Answer: The rat learns to press the lever repeatedly.
The Olds and Milner experiment demonstrated that rats learned lever pressing when the response produced electrical stimulation, showing that the stimulation could reinforce an instrumental behavior.
Key Takeaways
- The proposed actor-critic mapping associates the dorsal striatum with action selection and the ventral striatum with value learning and reward processing.
- The ventral striatum is proposed to send value information to dopamine neurons in the VTA and SNpc, where value and reward information contribute to activity corresponding to a temporal-difference error signal.
- Dopamine axons contact dendritic spines of medium spiny neurons, where dopamine may influence learning at cortical-to-striatal synapses.
- The Olds and Milner experiment showed that rats learned lever pressing when pressing produced electrical stimulation, supporting a reinforcing role for brain stimulation.
- Optogenetic studies support a reinforcing role for phasic dopamine activity, but the supplied material does not specify distinct effects for mice, rats, and fruit flies.