Reward Prediction Error
The TD error and phasic dopamine activity can show matching patterns under a defined set of assumptions.
Why Prediction Error Matters
A reward-prediction system should respond to more than the mere presence of a reward. It should also reflect whether the reward was expected, whether the expectation changed, and when the reward was expected to occur. The temporal-difference, or TD, model provides one computational account of this process. Its TD error can be compared with the short-lived, or phasic, activity of dopamine-producing neurons during classical conditioning.
The reward prediction error hypothesis proposes that phasic activity in dopamine-producing neurons can deliver an error between an old estimate and a new estimate of expected future reward to target areas throughout the brain.
Two Estimates One Error
The central change is a change in an estimate. First, a system has an old estimate of the reward expected in the future. Later, it forms a new estimate. The difference between the old and new estimates is treated as an error signal. In this account, the important event is not simply that a reward exists; it is that the expected-reward estimate changes.
A Simple Estimate Change
A system's old estimate of expected future reward is 5 units. After new information is available, its estimate is 8 units. Identify the direction and size of the error.
Identify the old estimate: The old estimate is 5 units.
Identify the new estimate: The new estimate is 8 units.
Compare them: The new estimate is 3 units higher than the old estimate.
The error is positive, with a size of 3 units. This illustrates an upward change in expected future reward; it does not by itself specify every detail of a neuron's physical activity.
Timing Inside the Representation
The complete serial compound, or CSC, representation gives the TD model separate internal signals for different moments after a cue. If a stimulus starts a different internal signal at each later time step, then each point between the cue and reward can be represented as a distinct state. Because the TD error depends on the state, the model can respond differently when a reward is expected shortly after a cue than when it is expected later.
Timing also matters when an expected reward fails to appear. The decrease in dopamine activity is described as occurring shortly after the time at which the reward was expected, rather than at an arbitrary point. The representation therefore supports both the timing of an expected reward and the timing of a missing expected reward.
How Learning Moves the Response
The clearest way to follow the learning process is to track where the phasic response appears. At the beginning, a neutral cue has little predictive value, so it does not produce a substantial phasic response. The rewarding event is the unexpected part of the trial. With continued learning, the cue gains predictive value and begins to elicit the phasic response. If another cue reliably occurs even earlier, the response can move again from the later predictive cue to the earlier one.
The response shifts in location because predictive value shifts. Initially, the reward carries the unexpected information. After learning, the cue carries information about the future reward, so the phasic response is associated with the cue instead. The reward itself is now predicted in that learned sequence.
Phasic Change and Background Activity
A phasic response is a short-lived change in activity, whereas background firing is the neuron's ongoing baseline activity. These should not be treated as the same quantity. In the comparison with dopamine-producing neurons, a negative TD error is represented as activity below the background firing rate. The hypothesis therefore concerns a change relative to background activity, not a claim that the neuron is simply either firing or silent.
Claims and Limits
The hypothesis makes a carefully scoped claim: the TD error and phasic dopamine activity can show matching patterns under a defined set of assumptions. It proposes that one function of phasic activity is to deliver the error between old and new expected-reward estimates to target areas throughout the brain. It does not claim that the TD error and dopamine activity are identical physical events, nor that the hypothesis concerns only whether a reward exists.
Treating the hypothesis as a claim that dopamine activity and the TD error are the same physical event.
The source describes a comparison of patterns under specific assumptions, not an identity between two physical events.
Fix:
Say that phasic dopamine activity is proposed to deliver an error whose patterns can match TD-error patterns.Defining prediction error as the presence or absence of reward alone.
The hypothesis tracks a change from an old estimate of expected future reward to a new estimate.
Fix:
Ask how the new expected-reward estimate differs from the old one.Ignoring timing.
The CSC representation supplies distinct internal states for different moments, allowing the TD error to respond differently to different expected timings.
Fix:
Represent the moments between cue and reward as separate states.Confusing below-background activity with no activity.
The negative response is described as activity below the background firing rate.
Fix:
Describe the response relative to the ongoing background rate.
When analyzing a prediction-error example, use this order: identify the old estimate, identify the new estimate, determine the direction of their difference, locate the relevant moment in the sequence, and then compare the predicted error pattern with phasic dopamine activity.
Check Your Understanding
A system has an old estimate of expected future reward of 9 units and a new estimate of 4 units. What is the direction of the prediction error? What kind of phasic dopamine pattern would the hypothesis associate with a negative error?
Hints
- Compare the new estimate with the old estimate.
- A negative error means the new estimate is lower than the old estimate.
- Use the background firing rate as the reference level.
What do you think happens?
A cue initially has little predictive value, but after repeated pairings it reliably predicts a reward. Where should the phasic response be expected to shift?
Reveal answer
Answer: From the reward toward the predictive cue
Initially, the reward is unexpected. With learning, the cue gains predictive value and begins to elicit the phasic response, while the reward becomes predicted in the learned sequence.
For the numerical prompt, the new estimate is 5 units lower than the old estimate, so the error is negative. In the hypothesis, a negative TD error corresponds to dopamine activity below its background firing rate, with the timing of the decrease related to when the reward was expected.
Key Takeaways
- A TD error is an error between an old and a new estimate of expected future reward.
- The CSC representation gives different moments in a cue-to-reward sequence distinct internal states, making the error sensitive to timing.
- As learning progresses, the phasic response can shift from an unexpected reward to an earlier predictive cue.
- A negative error is represented as dopamine activity below background firing, so a phasic decrease is not the same as silence.
- The reward prediction error hypothesis compares TD-error patterns with phasic dopamine activity under specific assumptions; it does not claim that the two are identical physical events.
Key Takeaways
- The reward prediction error hypothesis concerns a change between old and new estimates of expected future reward.
- The CSC representation preserves elapsed-time information by representing different moments as distinct states.
- Learning can move the phasic dopamine response from the reward to an earlier cue that predicts it.
- Positive, zero, and negative errors describe changes relative to an estimate and, in the comparison, correspond to increased, unchanged, or below-background phasic activity.
- The TD error and dopamine activity are compared as matching patterns under assumptions, not declared to be identical physical events.