Concepts / TD Error Concept

TD Error Concept

The hypothesis is about an error between old and new estimates of expected future reward.

  • Programming

A Changing Expectation

The TD error concept focuses on a change in an estimate. A system first has an old estimate of the reward expected in the future. It then forms a new estimate. The difference between the old and new estimates is treated as an error signal. The central idea is therefore not merely that a reward exists, but that the system's expected-reward estimate has changed.

new estimate formscompare with old estimateOld estimateExpected future rewardNew estimateExpected future rewardTD errorDifference betweenestimates
How do the old and new estimates differ, and how does their difference become the TD error?

Tracing the Two Estimates

Keep the sequence clear. First, there is an old estimate of expected future reward. Next, a new estimate becomes available. Finally, the two estimates are compared. The result of that comparison is the TD error. This sequence matters because the hypothesis is about an error between estimates, not simply about the presence or absence of a reward.

  1. Start with the old estimate of the reward expected in the future.
  2. Form a new estimate of the reward expected in the future.
  3. Compare the new estimate with the old estimate.
  4. Treat the difference between them as an error signal.
  5. Relate that error signal to the proposed function of phasic dopamine-producing neuron activity.
estimate changescompare estimatesdifference becomesOld estimateExpected future rewardNew estimateExpected future rewardDifferenceOld estimate and newestimateError signalTD error
What changes when a new expected-reward estimate differs from the old estimate?

A Numerical Difference

Comparing Two Reward Estimates

Suppose the old estimate of expected future reward is 5 units and the new estimate is 8 units. What is the difference between the estimates?

Identify the old estimate: The old estimate is 5 units.

Identify the new estimate: The new estimate is 8 units.

Compare them: The difference is found by comparing the new estimate with the old estimate: 8 minus 5 equals 3 units.

Interpret the result: The TD error in this illustration is a difference of 3 units. It represents the change between the two expected-reward estimates.

The difference between the old and new estimates is 3 units.

estimate changes8 minus 55Old estimate8New estimate3Difference
Given an old estimate of 5 and a new estimate of 8, what error value results from comparing them?

This example is intentionally simple. Its purpose is to separate the two quantities: 5 is the old estimate, 8 is the new estimate, and 3 is the difference between them. In a fuller reinforcement-learning treatment, the estimates and their update rules can be more elaborate, but the concept described here still begins with the difference between an old and a new expected-reward estimate.

Link to Phasic Dopamine

The TD error concept is connected to phasic activity in dopamine-producing neurons. The reward prediction error hypothesis proposes that one function of this phasic activity is to deliver the error signal to target areas throughout the brain. In this view, the activity is related to the difference between old and new estimates of expected future reward.

estimate changescompare estimatesproposed relationshipdeliver errorOld estimateExpected future rewardNew estimateExpected future rewardTD errorDifferencePhasic activityDopamine-producing neuronsTarget areasThroughout the brain
How is the TD error connected to phasic dopamine-producing neuron activity?

What the Hypothesis Does Not Say

  • Treating the hypothesis as a claim about reward alone.

    The hypothesis begins with an old estimate and a new estimate of expected future reward. The important quantity is their difference.

    Fix: Ask what changed in the expected-reward estimate, then identify the difference between the two estimates.

  • Confusing the old estimate with the new estimate.

    The error requires a comparison between two estimates that occupy different stages in the sequence.

    Fix: Label the starting value as the old estimate and the later value as the new estimate before comparing them.

  • Saying that the hypothesis proves phasic dopamine activity is identical to reward.

    The account describes phasic dopamine-producing neuron activity as proposed to deliver an error signal to target areas. That is different from saying the activity is the reward itself.

    Fix: State the relationship narrowly: the activity is proposed to deliver the error associated with the change in expected reward.

  • Turning a proposed function into an exhaustive explanation.

    The source says the TD error concept helps account for many features of phasic dopamine neuron activity and describes one proposed function. Those statements do not automatically cover every feature or function.

    Fix: Use qualified language such as proposed, helps account for, and one function.

Careful statementOverstatement to avoid
The hypothesis concerns the difference between old and new estimates of expected future reward.The hypothesis is simply about whether a reward exists.
Phasic dopamine-producing neuron activity is proposed to deliver the error signal to target areas.Phasic dopamine activity is identical to the reward.
The TD error concept helps account for many features of phasic dopamine neuron activity.The concept necessarily explains every feature and function of those neurons.

Check Your Understanding

EASY

A system's old estimate of expected future reward is 12 units. Its new estimate is 9 units. Identify the old estimate, the new estimate, and the difference between them. Then state how the result should be described in relation to the hypothesis.

Hints
  • Keep the two estimates separate before comparing them.
  • Compare 9 with 12 to find the difference.
  • Describe the result as an error between estimates, not simply as the reward.

What do you think happens?

If the old estimate is 12 units and the new estimate is 9 units, what is the numerical difference between the estimates?

  • 3 units
  • 9 units
  • 12 units
  • 21 units
Reveal answer

Answer: The difference in magnitude is 3 units; comparing the new estimate with the old estimate gives 9 minus 12, which is negative 3 units.

The key step is to compare two estimates rather than use either estimate by itself. The direction of the difference depends on which estimate is subtracted from which.

Key Takeaways

  1. TD error is the difference between an old estimate and a new estimate of expected future reward.
  2. The reward prediction error hypothesis focuses on a change in an estimate, not merely on the existence of a reward.
  3. Phasic dopamine-producing neuron activity is proposed to deliver this error signal to target areas throughout the brain.
  4. The TD error concept helps account for many features of phasic dopamine neuron activity, but the focused claim should not be expanded into an exhaustive explanation.
  5. A simple numerical illustration can identify the error by comparing the two reward estimates.

Key Takeaways

  • TD error is an error between old and new estimates of expected future reward.
  • The reward prediction error hypothesis treats this difference as an error signal.
  • Phasic dopamine-producing neuron activity is proposed to deliver the signal to target areas throughout the brain.
  • The hypothesis should not be overstated as saying that dopamine activity is simply the reward or that TD error explains every feature and function of dopamine-producing neurons.