Understanding Reward and Reinforcement Signals in Reinforcement Learning
Prediction error is the broad idea of measuring a mismatch between expectation and reality.
From Expectation to Error
Learning depends on comparing what was expected with what actually happens. The broad name for the mismatch between an expected signal and an actual signal or observation is prediction error. Reinforcement learning focuses this idea on reward, and then narrows it further when it compares long-term reward expectations across time.
An expectation is useful only when it can be compared with what actually happens. The prediction error is the resulting discrepancy. The signal being compared does not have to be a reward: it can concern any observation or event. This is the general case.
General and Reward-Specific Mismatches
A reward prediction error, or RPE, is not a separate kind of mismatch from the broad idea of prediction error. It is the reward-focused version. The expected signal is specifically an expected reward, and the actual signal is the reward that is received.
A Reward-Focused Comparison
A learner expects a reward and then receives a reward. Which kind of prediction error is being discussed?
Identify the expected signal: The expected signal concerns reward, so this is not merely an unspecified prediction error.
Identify the actual signal: The comparison uses the reward that was actually received.
Name the mismatch: The discrepancy between the expected reward and the received reward is a reward prediction error.
A reward prediction error is a prediction error whose comparison is specifically about expected and received reward.
Reading the Sign of an RPE
The sign of an RPE describes the direction of the mismatch. When the received reward is greater than expected, the RPE is positive. When the received reward is less than expected, the RPE is negative. When the received reward matches the expectation, there is no discrepancy between those two signals.
Suppose a learner expects a small reward but receives a larger one. The received reward exceeds the expectation, so the reward prediction error is positive. If the learner instead receives less than expected, the reward prediction error is negative. The important information is the direction of the mismatch, not merely the fact that a comparison occurred.
TD Error Across Time
A TD error narrows the reward-prediction idea further. It does not directly compare an expected reward with only the reward just received. Instead, it compares an earlier expectation of long-term reward with a current expectation of long-term reward.
Tracing a TD Error
A learner has an earlier expectation about long-term reward. After new information arrives, it has a current expectation about long-term reward. What comparison defines the TD error?
Start with the earlier expectation: This is the learner's earlier expectation of reward over the long term.
Form the current expectation: After the situation changes or new information is available, the learner has a current expectation of long-term reward.
Compare the two expectations: The TD error is the discrepancy between the earlier long-term reward expectation and the current long-term reward expectation.
A TD error is a special reward prediction error based on the change between earlier and current long-term reward expectations.
The key shift is what is being compared. An ordinary reward prediction error focuses on expected and received reward. A TD error compares expectations at different points in time: an earlier expectation and a current expectation of long-term reward. It therefore represents a discrepancy in long-term reward expectations.
Terminology in Neuroscience
The abbreviation RPE can be used broadly for reward prediction error. In neuroscience discussions, RPE may also be used to mean TD RPE, the temporal-difference version involving changing long-term reward expectations. In this chapter, the narrower case is called TD error so that the distinction is visible: RPE refers to the reward-focused comparison in general, while TD error refers to the comparison between earlier and current long-term reward expectations.
| Term | Comparison | Use in this chapter |
|---|---|---|
| Prediction error | An expected signal and an actual signal or observation | Broad category |
| Reward prediction error | An expected reward and a received reward | Reward-focused mismatch |
| TD error | An earlier long-term reward expectation and a current long-term reward expectation | Narrow temporal-difference case |
| RPE in some neuroscience discussions | May refer to the TD reward-prediction case | Terminology can overlap with TD RPE |
The terms overlap, but the comparison being made identifies the concept.
Mistakes to Avoid
Treating every prediction error as a reward prediction error.
Prediction error is the broad category. RPE is the reward-focused case.
Fix:
First identify the signal being compared. Use reward prediction error only when the comparison concerns expected and received reward.Defining TD error as only the difference between expected reward and received reward.
The TD comparison concerns an earlier and a current expectation of long-term reward.
Fix:
Ask whether the comparison involves long-term reward expectations at different points in time.Assuming the sign of an RPE describes the reward in isolation.
The sign indicates whether the received reward was greater than expected or less than expected.
Fix:
Compare the received reward with the expected reward before interpreting the sign.Assuming RPE always has exactly the same meaning across sources.
Neuroscience discussions may use RPE to refer to the TD case.
Fix:
Examine the comparison described by the source. In this chapter, use TD error for the long-term expectation comparison.
Check Your Understanding
Classify each comparison as a general prediction error, a reward prediction error, or a TD error. Then explain what determines the sign when the comparison is a reward prediction error: an expected signal and an actual observation; an expected reward and a received reward; or an earlier and current long-term reward expectation.
Hints
- Look first at whether the signal is specifically a reward.
- For TD error, look for two long-term reward expectations at different points in time.
- For the sign of an RPE, compare the received reward with the expected reward.
- Prediction error is the broad discrepancy between expectation and reality. A reward prediction error applies that comparison to expected and received reward. Its sign indicates whether the received reward was greater than or less than expected. A TD error is a narrower reward-prediction comparison between earlier and current long-term reward expectations. Although neuroscience discussions may use RPE for the TD case, this chapter calls that case TD error.
Key Takeaways
- Prediction error measures a mismatch between an expected signal and an actual signal or observation.
- Reward prediction error focuses the mismatch on expected and received reward.
- A positive RPE means the received reward was greater than expected; a negative RPE means it was less than expected.
- TD error is a special reward prediction error comparing earlier and current long-term reward expectations.
- Neuroscience sources may use RPE for the TD case, while this chapter uses TD error to make that narrower meaning explicit.