Reinforcement Learning and Reward Signals
Values forecast total reward over the future, while prediction errors describe mismatches between expectations and observations.
Forecasts Instead of Certainties
A reinforcement-learning agent often has several possible actions but cannot know the entire future. It therefore relies on value estimates. These estimates are forecasts of the total reward the agent may accumulate later, not descriptions of what is happening at this exact moment. When the reward or observation differs from the forecast, a prediction error occurs.
Values Attached to States and Actions
V estimates the long-term value of a state. In other words, it forecasts the total reward that may follow from being in that state. Q estimates the long-term value of an action. It therefore represents the expected future usefulness of taking a particular action in the current situation. Both are forecasts about future reward, but Q includes the consequence of choosing a particular action.
The useful comparison is not between a value and an immediate event. Instead, ask what each estimate forecasts over the future. V describes the future associated with a state, while Q distinguishes among the actions available in that state.
Choosing Among Available Actions
Suppose an agent is in one state and can select either Action A or Action B. The agent cannot inspect the entire future directly, so it compares its estimates of how worthwhile the available actions are. Those action-value estimates guide the choice toward the action forecast to lead to greater future reward.
Comparing two action forecasts
An agent is in one state and has two available actions. Its estimates suggest that Action A is more worthwhile over the future than Action B.
Start with the state: The agent identifies the current situation and the actions available from it.
Examine action values: The agent compares the long-term value forecast for Action A with the long-term value forecast for Action B.
Use the comparison: Because the estimate for Action A is more favorable, the value estimates guide the agent toward Action A.
The choice is guided by predicted future reward, not by a complete inspection of the future.
Two Kinds of Prediction Error
A reward prediction error compares an expected reward with the reward actually received. It focuses on the mismatch between an immediate reward expectation and the observed reward signal. A temporal-difference error, or TD error, compares expectations about long-term reward at different points. It concerns how a current or newer long-term forecast differs from an earlier expectation.
| Concept | What is compared? | What the mismatch describes |
|---|---|---|
| Reward prediction error | Expected reward and received reward | A difference between an expectation and an observed reward signal |
| Temporal-difference error | Earlier and current expectations about long-term reward | A difference between long-term forecasts at different points |
Reading the Direction of a Reward Error
To identify the direction of a reward prediction error, compare the observed reward with the expected reward. If the observed reward is greater than expected, the mismatch is positive. If the observed reward is less than expected, the mismatch is negative. If the observed and expected rewards are the same, there is no mismatch in this comparison.
Expectation compared with observation
Consider three generated cases in which an agent expects a reward and then receives an observed reward.
Case 1: The agent expects a reward of 5 and receives 8. The observation is greater than the expectation, so the reward prediction error is positive.
Case 2: The agent expects a reward of 5 and receives 5. The observation matches the expectation, so the reward prediction error is zero.
Case 3: The agent expects a reward of 5 and receives 2. The observation is less than the expectation, so the reward prediction error is negative.
The sign is determined by the direction of the mismatch between the observed reward and the expected reward.
Common Interpretation Mistakes
Treating V as the value of one particular action
V estimates the long-term value of a state. The action-specific estimate is Q.
Fix:
Use V when discussing a state and Q when discussing a particular action.Treating a value estimate as a description of the present
Values forecast total reward over the future; they are not descriptions of what is happening at the current instant.
Fix:
Interpret a value as a forecast about future reward.Calling every prediction error a reward prediction error
Reward prediction errors compare expected and received rewards, whereas TD errors compare expectations about long-term reward.
Fix:
Identify whether the comparison involves an observed reward or different long-term expectations.Assigning a sign without comparing expectation and observation
The direction of a reward prediction error depends on whether the observed reward is greater than, equal to, or less than the expectation.
Fix:
Compare the observed reward with the expected reward before deciding the sign.
Check Your Understanding
An agent expects a reward of 10 and receives a reward of 6. Is the reward prediction error positive, zero, or negative? Then explain whether this comparison is a reward prediction error or a temporal-difference error.
Hints
- First compare the observed reward with the expected reward.
- Then identify whether the comparison is between an expected reward and a received reward, or between long-term expectations.
An agent has one state and two available actions. Explain why Q estimates are more directly useful for comparing those actions than V alone.
Hints
- Recall what V estimates.
- Recall what additional information Q represents.
Key Takeaways
- Values forecast total future reward rather than merely describing the current moment.
- V estimates the long-term value of a state, while Q estimates the long-term value of a particular action.
- Value estimates guide action choices by allowing the agent to compare how worthwhile available actions are expected to be.
- A reward prediction error compares expected reward with received reward.
- A temporal-difference error compares earlier and current expectations about long-term reward.
- A reward prediction error is positive when the observed reward exceeds the expectation, zero when they match, and negative when the observed reward is smaller.
Key Takeaways
- State and action values are forecasts of future total reward.
- V describes a state's long-term value, while Q describes the long-term value of taking a particular action.
- Agents use value estimates to compare available actions and guide their choices.
- Reward prediction errors compare expected and received rewards, while TD errors compare expectations about long-term reward.
- The sign of a reward prediction error depends on whether the observation is greater than, equal to, or less than the expectation.