Concepts / Reinforcement Learning and Reward Signals

Reinforcement Learning and Reward Signals

Values forecast total reward over the future, while prediction errors describe mismatches between expectations and observations.

  • Programming

Forecasts Instead of Certainties

A reinforcement-learning agent often has several possible actions but cannot know the entire future. It therefore relies on value estimates. These estimates are forecasts of the total reward the agent may accumulate later, not descriptions of what is happening at this exact moment. When the reward or observation differs from the forecast, a prediction error occurs.

Values Attached to States and Actions

V estimates the long-term value of a state. In other words, it forecasts the total reward that may follow from being in that state. Q estimates the long-term value of an action. It therefore represents the expected future usefulness of taking a particular action in the current situation. Both are forecasts about future reward, but Q includes the consequence of choosing a particular action.

forecastforecastStatecurrent situationVlong-term value of thestateActionchoice in the stateQlong-term value of theaction
What does a state value predict, and what extra information does an action value provide?

The useful comparison is not between a value and an immediate event. Instead, ask what each estimate forecasts over the future. V describes the future associated with a state, while Q distinguishes among the actions available in that state.

Choosing Among Available Actions

Suppose an agent is in one state and can select either Action A or Action B. The agent cannot inspect the entire future directly, so it compares its estimates of how worthwhile the available actions are. Those action-value estimates guide the choice toward the action forecast to lead to greater future reward.

availableavailableestimateestimateguideCurrent stateAction Aaction-value estimateCompare valuesChosen actionforecast to be moreworthwhileAction Baction-value estimate
Given several possible actions in one state, how does the agent use predicted future rewards to guide its choice?

Comparing two action forecasts

An agent is in one state and has two available actions. Its estimates suggest that Action A is more worthwhile over the future than Action B.

Start with the state: The agent identifies the current situation and the actions available from it.

Examine action values: The agent compares the long-term value forecast for Action A with the long-term value forecast for Action B.

Use the comparison: Because the estimate for Action A is more favorable, the value estimates guide the agent toward Action A.

The choice is guided by predicted future reward, not by a complete inspection of the future.

Two Kinds of Prediction Error

A reward prediction error compares an expected reward with the reward actually received. It focuses on the mismatch between an immediate reward expectation and the observed reward signal. A temporal-difference error, or TD error, compares expectations about long-term reward at different points. It concerns how a current or newer long-term forecast differs from an earlier expectation.

comparecomparecomparecompareExpected rewardReward predictionerrormismatchEarlier expectationlong-term rewardTD errormismatchObserved rewardNewer forecastlong-term reward
How do the expected reward, observed reward, and updated long-term forecast differ between the two types of error?
ConceptWhat is compared?What the mismatch describes
Reward prediction errorExpected reward and received rewardA difference between an expectation and an observed reward signal
Temporal-difference errorEarlier and current expectations about long-term rewardA difference between long-term forecasts at different points

Reading the Direction of a Reward Error

To identify the direction of a reward prediction error, compare the observed reward with the expected reward. If the observed reward is greater than expected, the mismatch is positive. If the observed reward is less than expected, the mismatch is negative. If the observed and expected rewards are the same, there is no mismatch in this comparison.

Expectation compared with observation

Consider three generated cases in which an agent expects a reward and then receives an observed reward.

Case 1: The agent expects a reward of 5 and receives 8. The observation is greater than the expectation, so the reward prediction error is positive.

Case 2: The agent expects a reward of 5 and receives 5. The observation matches the expectation, so the reward prediction error is zero.

Case 3: The agent expects a reward of 5 and receives 2. The observation is less than the expectation, so the reward prediction error is negative.

The sign is determined by the direction of the mismatch between the observed reward and the expected reward.

Common Interpretation Mistakes

  • Treating V as the value of one particular action

    V estimates the long-term value of a state. The action-specific estimate is Q.

    Fix: Use V when discussing a state and Q when discussing a particular action.

  • Treating a value estimate as a description of the present

    Values forecast total reward over the future; they are not descriptions of what is happening at the current instant.

    Fix: Interpret a value as a forecast about future reward.

  • Calling every prediction error a reward prediction error

    Reward prediction errors compare expected and received rewards, whereas TD errors compare expectations about long-term reward.

    Fix: Identify whether the comparison involves an observed reward or different long-term expectations.

  • Assigning a sign without comparing expectation and observation

    The direction of a reward prediction error depends on whether the observed reward is greater than, equal to, or less than the expectation.

    Fix: Compare the observed reward with the expected reward before deciding the sign.

Check Your Understanding

EASY

An agent expects a reward of 10 and receives a reward of 6. Is the reward prediction error positive, zero, or negative? Then explain whether this comparison is a reward prediction error or a temporal-difference error.

Hints
  • First compare the observed reward with the expected reward.
  • Then identify whether the comparison is between an expected reward and a received reward, or between long-term expectations.
MEDIUM

An agent has one state and two available actions. Explain why Q estimates are more directly useful for comparing those actions than V alone.

Hints
  • Recall what V estimates.
  • Recall what additional information Q represents.

Key Takeaways

  1. Values forecast total future reward rather than merely describing the current moment.
  2. V estimates the long-term value of a state, while Q estimates the long-term value of a particular action.
  3. Value estimates guide action choices by allowing the agent to compare how worthwhile available actions are expected to be.
  4. A reward prediction error compares expected reward with received reward.
  5. A temporal-difference error compares earlier and current expectations about long-term reward.
  6. A reward prediction error is positive when the observed reward exceeds the expectation, zero when they match, and negative when the observed reward is smaller.

Key Takeaways

  • State and action values are forecasts of future total reward.
  • V describes a state's long-term value, while Q describes the long-term value of taking a particular action.
  • Agents use value estimates to compare available actions and guide their choices.
  • Reward prediction errors compare expected and received rewards, while TD errors compare expectations about long-term reward.
  • The sign of a reward prediction error depends on whether the observation is greater than, equal to, or less than the expectation.