Value Functions and Prediction Errors
An RPE is a difference signal, not the reward itself.
When Outcomes Defy Expectations
Reinforcement learning does not consider only whether a reward occurs. It also considers how the actual reward compares with the reward that was expected. That comparison produces a reward prediction error, or RPE. An RPE is therefore a difference signal: it represents a mismatch between expectation and outcome rather than the reward itself.
What do you think happens?
A learner expects a reward of 5 units and receives 8 units. What kind of reward prediction error should result?
Reveal answer
Answer: A positive prediction error because the outcome is better than expected
The outcome exceeds the expectation. The prediction error records that mismatch; it is not the reward of 8 units itself.
Reward Versus Mismatch
The reward is the outcome that occurs. The reward prediction error is a separate signal about the relationship between that outcome and the expectation that preceded it. If the actual reward is better than expected, the mismatch is positive. If it matches the expectation, the mismatch is zero. If it is worse than expected, the mismatch is negative. These descriptions concern the comparison, not the reward's identity.
| Item | What it represents | Question it answers |
|---|---|---|
| Reward | The actual outcome | What reward occurred? |
| Expected reward | The outcome anticipated by the learner or model | What reward was expected? |
| Reward prediction error | The difference between expectation and outcome | How did the outcome compare with expectation? |
Predictions Across Time
A temporal-difference error, or TD error, is a temporal-difference form of error signal. The temporal part matters because the comparison is considered across successive moments: a prediction at one moment is related to what is encountered at that moment and to the value predicted at another moment. In this framework, an error is not merely a label for an unexpected reward; it is part of a sequence of predictions and comparisons over time.
The diagram is a conceptual time sequence, not a numerical learning rule. Its key point is that TD error is the temporal-difference form of an error signal. The source makes a more specific neuroscience connection: phasic activity of dopamine-producing neurons is described as conveying TD errors.
Dopamine as a Candidate Error Signal
The reward prediction error hypothesis connects a computational idea from reinforcement learning with an observed biological signal. Reinforcement-learning theory defines an error signal based on expected and actual rewards. Neuroscience investigates activity in dopamine-producing neurons. The hypothesis proposes that dopamine neuron activity signals reward prediction errors, and that phasic activity of dopamine-producing neurons conveys temporal-difference errors.
Value as Future Expectation
A value function belongs to the prediction side of this discussion: it represents the expected future reward associated with a situation or state. The important contrast is between a prediction and the error produced when the later outcome differs from that prediction. A value estimate is therefore not identical to the reward that eventually occurs, and neither one is identical to the error signal describing their mismatch.
This flow separates three roles. The situation is what the learner or model is evaluating. The value function supplies an expectation about future reward. The prediction error records how the eventual outcome compares with that expectation. Keeping these roles separate prevents the value prediction or reward from being mistaken for the error signal.
A Three-Outcome Trace
Comparing One Expectation with Three Outcomes
Suppose an expectation is 5 units. Consider outcomes of 8 units, 5 units, and 2 units. Classify the prediction error in each case.
Outcome of 8: The actual outcome is greater than the expectation, so the mismatch is positive. In ordinary arithmetic language, 8 is 3 units above 5.
Outcome of 5: The actual outcome matches the expectation, so there is no mismatch between the two values.
Outcome of 2: The actual outcome is less than the expectation, so the mismatch is negative. In ordinary arithmetic language, 2 is 3 units below 5.
Interpretation: The three outcomes are rewards. The positive, zero, and negative descriptions refer to the prediction errors produced by comparing each reward with the expectation.
The same expected reward can be followed by a positive, zero, or negative prediction error, depending on the actual outcome.
| Expected reward | Actual reward | Prediction-error direction |
|---|---|---|
| 5 | 8 | Positive |
| 5 | 5 | Zero |
| 5 | 2 | Negative |
The direction of the RPE depends on the comparison between expected and actual reward.
Reading Neural Results Carefully
When examining a neuroscience result, follow the full reasoning chain. First identify the neural activity that was observed. Next compare its pattern with theoretical signal definitions. Then ask whether the activity resembles a reward signal, a value signal, a prediction error, a reinforcement signal, or another kind of signal. Only after that comparison should dopamine activity be discussed as evidence relevant to the reward prediction error hypothesis.
- Observe the neural activity.
- Compare the activity with the definitions of possible computational signals.
- Ask whether the activity resembles an RPE or a TD error.
- State the dopamine connection as a proposed interpretation rather than treating the label as automatic.
- Preserve the distinction between the reward, the expectation, and the mismatch signal.
Calling the reward itself an RPE.
The reward is 8; the RPE is the signal about the difference between 8 and 5.
Fix:
Name the outcome and the mismatch separately.Treating every dopamine response as automatically proving an RPE.
The source says that neural activity must be compared with several possible interpretations.
Fix:
Describe dopamine as a proposed or candidate signal for RPEs, then explain why the observed pattern resembles that theoretical signal.Using TD error and RPE as if the terms had exactly the same scope.
An RPE is the broad computational idea of a difference between expected and actual reward, while TD error is a type of error signal used in temporal-difference learning.
Fix:
State the broad RPE idea first, then identify TD error as the temporal-difference form.Ignoring the time dimension in TD error.
The temporal-difference description concerns error signals across time.
Fix:
Explain which predictions and outcomes are being related across successive moments.
Check Your Understanding
A model expects a reward of 10 units. It receives 10 units in one situation and 6 units in another. Explain what differs between the two cases. Then explain why neither the reward nor the prediction error should automatically be identified with dopamine activity.
Hints
- Compare the actual outcome with the expectation in each case.
- Use the words zero mismatch and negative mismatch where appropriate.
- Remember that the dopamine connection is described as a hypothesis linking neural activity with computational signals.
A strong answer should say that the outcome of 10 matches the expectation, whereas the outcome of 6 is worse than expected and therefore produces a negative mismatch. It should also preserve the interpretive chain: neural activity is observed, compared with theoretical signal definitions, and then considered in relation to the hypothesis that dopamine signals RPEs and that phasic dopamine-producing-neuron activity conveys TD errors.
Key Takeaways
- A reward prediction error is a difference signal based on the comparison between an expected reward and an actual reward.
- The reward is the outcome; the RPE is the signal about the mismatch between the outcome and the expectation.
- A temporal-difference error is a temporal-difference form of error signal.
- The reward prediction error hypothesis proposes that dopamine neuron activity signals RPEs.
- Phasic activity of dopamine-producing neurons is described as conveying TD errors, but neural activity should be compared with alternative interpretations before it is labeled an RPE.
Key Takeaways
- An RPE describes the difference between expected and actual reward, not the reward itself.
- Positive, zero, and negative RPEs correspond to outcomes that are better than, equal to, or worse than expected.
- TD error is a temporal-difference form of error signal.
- The dopamine connection is a hypothesis proposing that dopamine activity signals RPEs and that phasic dopamine-producing-neuron activity conveys TD errors.
- Interpret neural activity through a comparison with theoretical signal definitions rather than assigning the RPE label automatically.