Concepts / Values and Prediction Errors

Values and Prediction Errors

Prediction error is the broad idea of measuring a mismatch between expectation and reality.

  • Programming

When Expectations Meet Events

An expectation is useful only when it can be compared with what actually happens. If an agent expects one signal but receives another, the mismatch between the expectation and the observation is a prediction error. The idea is broad: the signal may concern any observation or event. Reinforcement learning makes the idea more specific by asking what happens when that signal is a reward.

comparecomparecomparecompareExpected signalany signal or eventExpected rewardreward forecastActual observationwhat happensReceived rewardreward observedPrediction errormismatchReward predictionerrorreward mismatch
What changes when the signal being compared is specifically a reward?

Reading the Error Sign

A reward prediction error, or RPE, focuses the comparison on expected reward and received reward. Its sign describes the direction of the mismatch. When the received reward is greater than expected, the RPE is positive. When the received reward is lower than expected, the RPE is negative. If the received reward matches the expectation, there is no mismatch between those two reward signals.

received reward is greaterreceived reward matchesreceived reward is lowerExpected rewardreference pointGreater rewardpositive RPEMatching rewardno mismatchLower rewardnegative RPE
Given an expected reward and an observed reward, when does the mismatch point upward, downward, or disappear?

Classifying Three Reward Outcomes

Compare an expected reward with the reward that is actually received.

Expected 5, received 8: The received reward is greater than expected, so the reward prediction error is positive.

Expected 5, received 5: The received reward matches the expectation, so there is no mismatch between the expected and received reward.

Expected 5, received 2: The received reward is lower than expected, so the reward prediction error is negative.

Greater-than-expected reward gives a positive RPE; matching reward gives no mismatch; lower-than-expected reward gives a negative RPE.

Values as Future Forecasts

Values are forecasts, not descriptions of the reward happening at this exact moment. A state value, written as V in the source terminology, estimates the long-term value of a state. An action value, written as Q, estimates the long-term value of an action. In this way, values forecast the total reward the agent may accumulate later.

maps tomaps toforecastsforecastsStatesState-action pairs, aFuture rewardtotal reward forecastVlong-term value of a stateQlong-term value of anaction
What does each kind of value estimate about future reward?
EstimateWhat it is aboutWhat it predicts
VA stateThe state's long-term value
QAn actionThe action's long-term value

From Values to Actions

An agent cannot know the entire future when it must choose what to do. It therefore relies on estimates of how worthwhile the current state is and how worthwhile each available action may be. Comparing action values gives the agent information it can use while choosing among those actions. The estimates guide the choice; they are forecasts about later reward rather than guarantees of what will happen.

presentsevaluateguideCurrent statewhere the agent isAvailable actionspossible choicesAction valueslong-term forecastsAction choiceguided by estimates
How does an agent use estimates of available actions to guide its next choice?

Imagine an agent facing two available actions. It cannot inspect the entire future directly, so it uses its estimates of how worthwhile the actions are over the long term. Those estimates help it decide what to do next. Later observations or rewards can differ from the forecasts, producing prediction errors.

The Temporal-Difference Step

A temporal-difference error, or TD error, is a more specific kind of reward prediction error. A usual reward prediction error compares an expected reward with the reward just received. A TD error instead compares two expectations about long-term reward: an earlier expectation and a current expectation. The immediate reward and the change between those long-term expectations are combined in the TD comparison.

comparecomparecombinecombineEarlier expectationlong-term rewardExpectationdiscrepancycurrent versus earlierTD errorspecial reward predictionerrorImmediate rewardreceived signalCurrent expectationlong-term reward
How are an immediate reward and the difference between current and earlier long-term reward expectations combined in a TD error?
Error typeSignals being comparedWhat makes it distinctive
Prediction errorAn expected signal and an actual signal or observationBroadest category
Reward prediction errorExpected reward and received rewardFocuses on reward
TD errorCurrent and earlier long-term reward expectations, together with the immediate rewardA more specific reward prediction error

RPE and TD Error Language

The labels can overlap. A TD error is a special kind of reward prediction error, so neuroscience discussions may use RPE as a broad label that includes what this chapter calls a TD RPE. This chapter uses TD error when it specifically means the comparison involving current and earlier long-term reward expectations. Keeping the narrower term visible helps distinguish that comparison from a direct expected-reward versus received-reward comparison.

Common Classification Mistakes

  • Treating every prediction error as a reward prediction error.

    Prediction error is the broad category. An RPE is the narrower case in which the expected and observed signals concern reward.

    Fix: First identify the signal. Use RPE only when the comparison concerns expected and received reward.

  • Describing a value as the reward happening now.

    Values forecast total or long-term reward over the future.

    Fix: Remember that V estimates the long-term value of a state and Q estimates the long-term value of an action.

  • Confusing a direct RPE comparison with a TD comparison.

    A TD error changes the comparison: it concerns earlier and current long-term reward expectations, along with the immediate reward.

    Fix: Name the error as a TD error when the comparison involves changing long-term expectations.

  • Reversing the sign of the reward prediction error.

    The sign indicates whether received reward was greater or lower than expected.

    Fix: Greater than expected is positive; lower than expected is negative.

Check Your Understanding

EASY

An agent expects a reward of 10 and receives a reward of 6. Is the reward prediction error positive or negative? Is this information by itself enough to describe a TD error?

Hints
  • Compare the received reward with the expected reward.
  • For the second question, ask whether the comparison includes current and earlier long-term reward expectations.

What do you think happens?

An agent expects 10 units of reward and receives 6. What should you predict about the direct reward prediction error?

  • Positive
  • Negative
  • No mismatch
Reveal answer

Answer: Negative

The received reward is lower than the expected reward. This identifies the direction of the direct reward prediction error; it does not by itself establish a TD error because no earlier and current long-term expectations were specified.

Key Takeaways

  1. A prediction error is a mismatch between an expected signal and an actual signal or observation.
  2. A reward prediction error focuses that mismatch on expected and received reward.
  3. A received reward greater than expected produces a positive RPE; a lower reward produces a negative RPE.
  4. A TD error is a more specific reward prediction error involving current and earlier long-term reward expectations together with the immediate reward.
  5. V estimates the long-term value of a state, while Q estimates the long-term value of an action; these estimates help guide action choices.

Key Takeaways

  • Prediction error is the broad mismatch between expectation and reality.
  • Reward prediction error narrows the comparison to expected and received reward.
  • The sign of an RPE shows whether the received reward was greater or lower than expected.
  • TD error is the specific case involving current and earlier long-term reward expectations.
  • State and action values forecast future reward and give an agent information for choosing among actions.