TD Errors
The reward prediction error hypothesis links dopamine neuron activity with reward prediction errors.
From Rewards to Learning Signals
Reinforcement learning theory describes signals that help an agent learn from the consequences of its behavior. Neuroscientists have searched for brain activity that could play a comparable role. The reward prediction error hypothesis proposes that dopamine neuron activity represents one such signal: a reward prediction error. The hypothesis became influential because the behavior of TD errors in reinforcement learning theory shows striking parallels with the phasic activity of dopamine neurons.
A TD error is not simply a description of whether a reward occurred. It is a learning quantity that compares an existing estimate with a better estimate formed after new information arrives. This makes TD errors relevant to dopamine research: both are studied as signals that may indicate how learning should respond to what happens over time.
Dopamine Prediction Errors
The reward prediction error hypothesis links dopamine neuron activity with reward prediction errors. Its central evidence is the parallel between TD-error behavior and phasic dopamine activity. Phasic dopamine activity functions as a reinforcement signal that reaches multiple brain areas and supports learning.
Wolfram Schultz and colleagues conducted influential experiments in the late 1980s and 1990s. Those experiments revealed parallels between dopamine neuron activity and TD errors, helping lead scientists to the reward prediction error hypothesis. The important contribution was the comparison between a computational description and experimental observations, not the name of dopamine by itself.
The One-Step Comparison
The TD error is the difference between the current estimate V(S_t) and the one-step estimate R_t+1 + γV(S_t+1). The one-step estimate is formed from the next reward and the discounted estimate of the next state.
δ_t = [R_t+1 + γV(S_t+1)] − V(S_t)The expression has two sides to compare. V(S_t) is the current estimate for the state at time t. R_t+1 + γV(S_t+1) is a one-step-better estimate because it uses information from the next transition: the reward received and the estimated value of the next state. The TD error records the difference between that newer one-step estimate and the old estimate.
Reading a TD Error Symbolically
An agent has a current estimate V(S_t), then observes the next reward R_t+1 and the next state S_t+1. Which information belongs in the one-step estimate?
Start with the current estimate: The old estimate is V(S_t), the value assigned to the state at time t.
Add the next reward: The transition supplies R_t+1, the reward received at the next time step.
Include the next-state estimate: The value of S_t+1 is discounted and contributes γV(S_t+1).
Compare the estimates: The TD error is the one-step estimate R_t+1 + γV(S_t+1) minus the old estimate V(S_t).
The sample backup target is R_t+1 + γV(S_t+1), and the TD error is its difference from V(S_t).
Why the Error Arrives Later
The label δ_t refers to the error in the estimate V(S_t), but calculating it requires R_t+1 and S_t+1. Those are next-step observations. Therefore, the TD error associated with time t becomes available at time t+1.
The subscript t identifies which estimate is being evaluated. It does not mean that every quantity needed for the calculation was already known at time t.
Building the Backup Target
After the transition from S_t to S_t+1, two pieces of new information are available: the next reward R_t+1 and the next state S_t+1. The reward is included directly. The estimate of the next state is multiplied by γ, the discount factor, and then included. Together they form the sample backup target R_t+1 + γV(S_t+1).
This target is called a sample backup target because it uses the particular reward and next state observed in one transition. The TD error then compares this target with the estimate that existed for S_t.
What do you think happens?
Which quantity must be observed before δ_t can be calculated?
Reveal answer
Answer: The next reward R_t+1 and next state S_t+1
The one-step estimate requires both R_t+1 and γV(S_t+1). Without the next transition information, the comparison cannot be completed.
Three Learning Signals
The sign of the TD error describes how the one-step estimate compares with the current estimate. A positive error means the one-step estimate is greater than V(S_t); a negative error means it is lower; and a zero error means the two estimates match. As a reinforcement signal, this difference can indicate whether the current estimate should move upward, downward, or not change in response to the new transition.
This is why phasic dopamine activity is discussed as a possible reinforcement signal. The hypothesis connects changes in dopamine neuron activity with a computational difference that tells a learning system how new consequences compare with its prior prediction.
TD and Monte Carlo Errors
TD error and Monte Carlo error use different references. TD error uses a one-step sample backup: the next reward plus the discounted estimate of the next state. Monte Carlo error is based on the actual return. The source gives an important relationship between them: the Monte Carlo error can be written as a sum of TD errors.
| Error | Reference | Time perspective |
|---|---|---|
| TD error | One-step sample backup | Uses the next reward and next-state estimate |
| Monte Carlo error | Actual return | Can be represented as a sum of TD errors |
Monte Carlo error = sum of TD errorsCommon Interpretation Mistakes
Treating a TD error as the reward itself
The TD error compares a current estimate with a one-step estimate that includes both the next reward and the discounted next-state estimate.
Fix:
Use δ_t = [R_t+1 + γV(S_t+1)] − V(S_t).Assuming δ_t is available before the next transition
The calculation requires R_t+1 and S_t+1, which are next-step observations.
Fix:
Associate δ_t with the estimate at time t, but recognize that it becomes available at time t+1.Confusing the TD target with the TD error
That expression is the one-step estimate or sample backup target. The error is its difference from V(S_t).
Fix:
First form the target, then subtract the current estimate.Treating Monte Carlo error and TD error as the same comparison
TD error uses a one-step sample backup, whereas Monte Carlo error is based on the actual return.
Fix:
Remember that the Monte Carlo error can be represented as a sum of TD errors.
Check Your Understanding
Explain, in your own words, why the error labeled δ_t is not available until time t+1. Then identify the two quantities that must be compared to calculate it.
Hints
- Start with the information known at time t.
- Identify what the transition reveals at time t+1.
- State the current estimate and the one-step estimate separately.
Compare TD error with Monte Carlo error. Your answer should name the reference used by each error and explain how a sequence of TD errors relates to the Monte Carlo error.
Hints
- TD error uses a one-step sample backup.
- Monte Carlo error uses the actual return.
- The source describes the Monte Carlo error as a sum of TD errors.
Key Takeaways
- The reward prediction error hypothesis proposes that dopamine neuron activity represents a reward prediction error.
- TD errors compare the current estimate V(S_t) with the one-step estimate R_t+1 + γV(S_t+1).
- The error labeled δ_t becomes available at time t+1 because the next reward and next state are required.
- Phasic dopamine activity is studied as a possible reinforcement signal because its activity shows parallels with TD-error behavior.
- Monte Carlo error uses the actual return and can be written as a sum of TD errors.
Key Takeaways
- A TD error measures the difference between an existing value estimate and a one-step-better estimate.
- The one-step estimate combines the next reward with the discounted estimate of the next state.
- Although δ_t concerns the estimate at time t, it becomes available at time t+1 because next-step information is needed.
- The reward prediction error hypothesis connects this reinforcement-learning signal with phasic dopamine activity.
- Monte Carlo error is based on the actual return and can be represented as a sum of TD errors.