Values and Prediction Errors
Prediction error is the broad idea of measuring a mismatch between expectation and reality.
When Expectations Meet Events
An expectation is useful only when it can be compared with what actually happens. If an agent expects one signal but receives another, the mismatch between the expectation and the observation is a prediction error. The idea is broad: the signal may concern any observation or event. Reinforcement learning makes the idea more specific by asking what happens when that signal is a reward.
Reading the Error Sign
A reward prediction error, or RPE, focuses the comparison on expected reward and received reward. Its sign describes the direction of the mismatch. When the received reward is greater than expected, the RPE is positive. When the received reward is lower than expected, the RPE is negative. If the received reward matches the expectation, there is no mismatch between those two reward signals.
Classifying Three Reward Outcomes
Compare an expected reward with the reward that is actually received.
Expected 5, received 8: The received reward is greater than expected, so the reward prediction error is positive.
Expected 5, received 5: The received reward matches the expectation, so there is no mismatch between the expected and received reward.
Expected 5, received 2: The received reward is lower than expected, so the reward prediction error is negative.
Greater-than-expected reward gives a positive RPE; matching reward gives no mismatch; lower-than-expected reward gives a negative RPE.
Values as Future Forecasts
Values are forecasts, not descriptions of the reward happening at this exact moment. A state value, written as V in the source terminology, estimates the long-term value of a state. An action value, written as Q, estimates the long-term value of an action. In this way, values forecast the total reward the agent may accumulate later.
| Estimate | What it is about | What it predicts |
|---|---|---|
| V | A state | The state's long-term value |
| Q | An action | The action's long-term value |
From Values to Actions
An agent cannot know the entire future when it must choose what to do. It therefore relies on estimates of how worthwhile the current state is and how worthwhile each available action may be. Comparing action values gives the agent information it can use while choosing among those actions. The estimates guide the choice; they are forecasts about later reward rather than guarantees of what will happen.
Imagine an agent facing two available actions. It cannot inspect the entire future directly, so it uses its estimates of how worthwhile the actions are over the long term. Those estimates help it decide what to do next. Later observations or rewards can differ from the forecasts, producing prediction errors.
The Temporal-Difference Step
A temporal-difference error, or TD error, is a more specific kind of reward prediction error. A usual reward prediction error compares an expected reward with the reward just received. A TD error instead compares two expectations about long-term reward: an earlier expectation and a current expectation. The immediate reward and the change between those long-term expectations are combined in the TD comparison.
| Error type | Signals being compared | What makes it distinctive |
|---|---|---|
| Prediction error | An expected signal and an actual signal or observation | Broadest category |
| Reward prediction error | Expected reward and received reward | Focuses on reward |
| TD error | Current and earlier long-term reward expectations, together with the immediate reward | A more specific reward prediction error |
RPE and TD Error Language
The labels can overlap. A TD error is a special kind of reward prediction error, so neuroscience discussions may use RPE as a broad label that includes what this chapter calls a TD RPE. This chapter uses TD error when it specifically means the comparison involving current and earlier long-term reward expectations. Keeping the narrower term visible helps distinguish that comparison from a direct expected-reward versus received-reward comparison.
Common Classification Mistakes
Treating every prediction error as a reward prediction error.
Prediction error is the broad category. An RPE is the narrower case in which the expected and observed signals concern reward.
Fix:
First identify the signal. Use RPE only when the comparison concerns expected and received reward.Describing a value as the reward happening now.
Values forecast total or long-term reward over the future.
Fix:
Remember that V estimates the long-term value of a state and Q estimates the long-term value of an action.Confusing a direct RPE comparison with a TD comparison.
A TD error changes the comparison: it concerns earlier and current long-term reward expectations, along with the immediate reward.
Fix:
Name the error as a TD error when the comparison involves changing long-term expectations.Reversing the sign of the reward prediction error.
The sign indicates whether received reward was greater or lower than expected.
Fix:
Greater than expected is positive; lower than expected is negative.
Check Your Understanding
An agent expects a reward of 10 and receives a reward of 6. Is the reward prediction error positive or negative? Is this information by itself enough to describe a TD error?
Hints
- Compare the received reward with the expected reward.
- For the second question, ask whether the comparison includes current and earlier long-term reward expectations.
What do you think happens?
An agent expects 10 units of reward and receives 6. What should you predict about the direct reward prediction error?
Reveal answer
Answer: Negative
The received reward is lower than the expected reward. This identifies the direction of the direct reward prediction error; it does not by itself establish a TD error because no earlier and current long-term expectations were specified.
Key Takeaways
- A prediction error is a mismatch between an expected signal and an actual signal or observation.
- A reward prediction error focuses that mismatch on expected and received reward.
- A received reward greater than expected produces a positive RPE; a lower reward produces a negative RPE.
- A TD error is a more specific reward prediction error involving current and earlier long-term reward expectations together with the immediate reward.
- V estimates the long-term value of a state, while Q estimates the long-term value of an action; these estimates help guide action choices.
Key Takeaways
- Prediction error is the broad mismatch between expectation and reality.
- Reward prediction error narrows the comparison to expected and received reward.
- The sign of an RPE shows whether the received reward was greater or lower than expected.
- TD error is the specific case involving current and earlier long-term reward expectations.
- State and action values forecast future reward and give an agent information for choosing among actions.