Reward Prediction Error Hypothesis
Dopamine neuron activity can shift from a reward to an earlier stimulus that predicts the reward.
The Surprising Shift
A first interpretation of dopamine activity might be that dopamine neurons simply respond to food, to the movement that begins when an animal sees food, or to the reward itself. The experiments discussed in this topic point to a more specific explanation. As learning develops, phasic dopamine activity can move away from the later reward and toward an earlier stimulus that predicts when and where the reward will be available.
The central observation is not that every event produces the same dopamine response. It is that the timing of a response can change: a monitored dopamine response that initially occurs at reward delivery may later occur when a reliable predictive cue appears.
Following the Learning Sequence
The sequence can be understood as a movement through time. Before learning, the reward is the important event because it is not yet predicted. After repeated learning, an earlier cue provides information about the reward. The dopamine response can therefore appear at that cue instead of at the later reward. If an even earlier cue becomes informative, the response can move earlier again. As the response moves earlier, it can disappear from later predictors.
The Lever-Task Sequence
Trace where many monitored dopamine neurons responded as the monkeys learned the lever task.
Initial stage: At first, many dopamine neurons responded when apple juice was delivered.
After continued training: Many of those neurons stopped responding to the juice and instead responded when the trigger light was illuminated.
Earlier predictor: In the lever tasks, the response could move again to an earlier instruction cue.
Behavioral change: The monkeys pressed the lever faster as training continued.
The important pattern is a shift in response timing: from the reward, to the trigger light, and then to an earlier instruction cue.
Reading the Prediction Error
The reward prediction error hypothesis links dopamine neuron activity with reward prediction errors. Reinforcement learning describes a temporal-difference error as a quantity in a learning process that unfolds over time. The source presents the important teaching point as a correspondence: the behavior of TD errors and the pattern of phasic dopamine activity can be studied as parallel signals involved in learning.
A reward prediction error concerns the relationship between what is expected and what is received. Early in learning, the reward is not yet reliably predicted, so the reward event is associated with the important phasic response. Once a cue reliably predicts the reward, information about the reward has shifted earlier in the sequence. The response can then occur at the cue, while the later reward produces less of the monitored response.
Reward Versus Predictor
| Event in the learned sequence | What the experiments show |
|---|---|
| Unexpected reward | Many monitored dopamine neurons can respond when the reward is delivered. |
| Reliable predictive cue | The response can shift to the earlier cue that predicts the reward. |
| Later reward after learning | The response can diminish as the response moves to the earlier predictor. |
| Earlier instruction cue | In the lever tasks, the response could move to an even earlier cue. |
This distinction prevents a common misunderstanding. A response at the reward does not necessarily mean that dopamine neurons are responding only to the sensory features of food. A response at a cue does not mean that the cue is itself the final reward. The cue matters because learning has made it informative about the future reward.
In the food-bin experiment, after training, neurons responded to the sight and sound of the opening rather than to touching the food. This result illustrates the same distinction: the response was associated with an event that predicted access to food rather than simply with contact with the food itself.
Dopamine as a Learning Signal
Reinforcement learning theories describe signals that help an agent learn from the consequences of its behavior. The reward prediction error hypothesis proposes that phasic dopamine activity represents one such signal. Because the activity can occur at a predictive cue, it can support learning about the cue and the behavior that leads through the cue toward the reward.
Phasic dopamine activity functions as a reinforcement signal that reaches multiple brain areas and supports learning. The hypothesis is therefore about a possible computational role for the activity, not merely about a reaction to the sensory properties of food.
Treating dopamine activity as a response that must remain attached to the reward.
The experiments show that the response can move to an earlier stimulus that predicts the reward.
Fix:
Track the timing of the response across learning rather than focusing only on the reward event.Assuming the cue and the reward are the same event.
The trigger light is described as a stimulus that predicts the later reward.
Fix:
Keep the cue and the reward separate: the cue provides predictive information, while the reward is the later outcome.Claiming that every event in the task produces the same dopamine response.
The reported result concerns a change in when monitored dopamine neurons respond, with later responses able to diminish as an earlier predictor gains a response.
Fix:
Describe the shift in timing and the changing response at later predictors.
Schultz's Experimental Contribution
Wolfram Schultz and colleagues conducted influential experiments in the late 1980s and 1990s. Their work helped reveal a neural-computational parallel: the way TD errors behave in reinforcement learning theory resembles the way phasic dopamine activity changes as rewards become predictable.
The experiments were influential because they connected a measurable neural pattern with a computational idea. Across the food-bin and lever tasks, the response did not simply remain tied to food delivery. It shifted toward sights, sounds, lights, and instruction cues that predicted the reward. This parallel between TD-error behavior and phasic dopamine activity helped lead scientists to the reward prediction error hypothesis.
Check the Mechanism
A task begins with an unexpected reward. After training, a light reliably appears before the reward, and an instruction cue appears before the light. Describe the expected movement of the monitored phasic dopamine response across the sequence.
Hints
- Start by identifying the event that matters before the reward can be predicted.
- Then identify the earliest cue that has become informative.
- Remember that a response can diminish at a later event when it shifts to an earlier predictor.
What do you think happens?
After learning, which sequence best matches the reported shift in the lever tasks?
Reveal answer
Answer: The response shifts from apple juice to the trigger light and then to the earlier instruction cue.
As earlier stimuli become reliable predictors, the monitored phasic response can move earlier in the sequence, while later responses can diminish.
Key Takeaways
- Phasic dopamine activity can shift from an unexpected reward to an earlier stimulus that predicts the reward.
- In the lever tasks, the response moved from apple juice to a trigger light and then to an earlier instruction cue.
- As an earlier predictor gains the response, the response at a later reward or predictor can diminish.
- The reward prediction error hypothesis proposes that dopamine neuron activity represents a reinforcement-learning signal related to reward prediction errors.
- Schultz's experiments were influential because they revealed parallels between TD errors and phasic dopamine activity.