Concepts / Dopamine Neuron Activity

Dopamine Neuron Activity

The hypothesis is about an error between old and new estimates of expected future reward.

  • Programming

A Change in Expectation

The reward prediction error hypothesis begins with a change in an estimate. A system first has an old estimate of the reward expected in the future. It then forms a new estimate. The difference between the old and new estimates is treated as an error signal. The hypothesis proposes that one function of phasic activity in dopamine-producing neurons is to deliver this error to target areas throughout the brain.

Tracing the Two Estimates

new estimate formscomparecompareOld estimateExpected future rewardEstimate differenceError signalNew estimateUpdated expected futurereward
What changes when the system replaces an earlier expected-reward estimate with a later one?

The old estimate is the expectation held before the update. The new estimate is the expectation available afterward. The error is produced by comparing these two estimates. This makes the hypothesis about an update in expected future reward rather than about the mere presence of a reward.

From TD Error to Phasic Activity

estimate updatescomparecompareproposed signaldeliverOld estimateExpected future rewardNew estimateUpdated expectationTD errorDifference betweenestimatesPhasic activityDopamine-producing neuronsTarget areasThroughout the brain
How does a change between reward estimates connect to phasic dopamine neuron activity?

In reinforcement learning, the TD error concept helps account for many features of phasic dopamine neuron activity. In the hypothesis described here, the difference between the old and new expected-reward estimates is the error signal, and phasic dopamine-producing neuron activity is proposed to deliver that signal to target areas throughout the brain.

The claim is therefore a relationship among three ideas: an expected-reward estimate changes, the difference is treated as an error, and phasic dopamine-producing neuron activity is proposed to deliver that error. This is narrower than saying that dopamine activity is identical to reward itself.

A Numerical Estimate Update

Comparing an old estimate with a new estimate

Suppose an old expected-reward estimate is 4 units and a later estimate is 7 units. Using the generated convention that error equals new estimate minus old estimate, identify the difference and its direction.

Identify the old estimate: The earlier expected-reward estimate is 4 units.

Identify the new estimate: The later expected-reward estimate is 7 units.

Compare the estimates: Subtract the old estimate from the new estimate: 7 minus 4 equals 3 units.

Interpret the direction: Because the new estimate is larger than the old estimate, the difference is positive under this convention.

The estimate difference is positive 3 units. The numerical illustration shows how an updated expected-reward estimate can differ from the earlier estimate.

Old estimateNew estimateDifference using new minus oldInterpretation
47Positive 3The new estimate is higher
74Negative 3The new estimate is lower
550The estimates are unchanged

What the Hypothesis Does Not Say

  • Treating the hypothesis as the claim that dopamine activity simply equals reward.

    The hypothesis focuses on the difference between an old and a new estimate of expected future reward, not simply on the fact that a reward exists.

    Fix: Ask what changed in the expected-reward estimate before asking what error signal is proposed.

  • Treating dopamine activity as identical to pleasure.

    The source describes phasic dopamine-producing neuron activity as proposed delivery of an error signal to target areas.

    Fix: State the narrower claim: phasic activity is proposed to deliver the difference between old and new expected-reward estimates.

  • Ignoring the old estimate.

    Without the old estimate, there is no comparison from which to identify the error described by the hypothesis.

    Fix: Write down both estimates and compare them.

  • Treating the hypothesis as a complete definition of motivation.

    The source makes a more limited proposal about one function of phasic activity and its delivery of an error signal.

    Fix: Avoid extending the claim beyond the proposed relationship among estimates, error, and phasic dopamine-producing neuron activity.

When explaining the hypothesis, use a precise sequence: name the old expected-reward estimate, name the new estimate, identify their difference as the error, and then describe phasic dopamine-producing neuron activity as the proposed delivery of that error to target areas.

A Careful Statement

EASY

An old expected-reward estimate is 10 units and a new estimate is 6 units. Using new estimate minus old estimate, identify the difference and explain what has changed.

Hints
  • Start with the later estimate and subtract the earlier estimate.
  • Compare the signs of the two estimates only after calculating the difference.
  • Use the words lower, higher, or unchanged to describe the update.

What do you think happens?

What is the difference when the new estimate is 6 units and the old estimate is 10 units, using new estimate minus old estimate?

  • Positive 4 units
  • Negative 4 units
  • Zero units
Reveal answer

Answer: Negative 4 units

The calculation is 6 minus 10, which equals negative 4. The new estimate is lower than the old estimate.

A careful summary of the hypothesis is: a system updates an estimate of expected future reward; the difference between the old and new estimates is treated as an error; and one proposed function of phasic activity in dopamine-producing neurons is to deliver that error to target areas throughout the brain. The TD error concept from reinforcement learning helps account for many features of this phasic activity.

Key Takeaways

  • The reward prediction error hypothesis concerns a difference between old and new estimates of expected future reward.
  • The old estimate is the expectation held before an update, while the new estimate is the expectation formed afterward.
  • The TD error concept helps account for many features of phasic dopamine neuron activity.
  • Phasic activity in dopamine-producing neurons is proposed to deliver the error signal to target areas throughout the brain.
  • The hypothesis should not be casually restated as dopamine simply being reward, pleasure, or motivation.