Understanding Reward Signals
A reward is a scalar number passed from the environment to the agent at each time step.
Why One Number Matters
In reinforcement learning, an agent can be described informally as having a purpose: it should achieve some desirable outcome. A reward signal turns that purpose into something measurable. At each time step, the environment passes the agent a scalar number called a reward. The agent is then understood to seek more reward over the long run, rather than judging success from only one number at one moment.
A reward is not a paragraph describing what the agent should do. It is a scalar number supplied by the environment at each time step.
From One Reward to a Sequence
The reward received at the current time step is an immediate reward. It tells the agent about one point in the interaction. Reinforcement learning, however, is concerned with cumulative reward over time. That means the rewards from multiple time steps are considered together as a sequence, and the sequence is judged by its total rather than by only its first or next reward.
Adding a Short Reward Sequence
An outcome produces rewards of 2, 3, and 1 at three successive time steps. What is its cumulative reward?
List the rewards: The sequence contains three scalar rewards: 2, 3, and 1.
Combine the time steps: Consider the rewards together instead of judging only the first or next reward.
Calculate the total: Adding 2, 3, and 1 gives a cumulative reward of 6.
The outcome has a cumulative reward of 6.
The Reward Hypothesis
The reward hypothesis formalizes the agent's purpose as maximizing the expected value of cumulative reward. This connects the three central ideas: the environment supplies a simple number at each time step, those numbers are accumulated across time, and the agent's goal is expressed through the expected value of that accumulated result.
Comparing Complete Outcomes
To apply the reward hypothesis, compare complete reward sequences and their totals. A sequence that begins with a larger immediate reward does not necessarily produce the larger cumulative reward. The relevant comparison is the total across the time steps being considered.
| Outcome | Reward sequence | Cumulative reward |
|---|---|---|
| A | 4, 0, 1 | 5 |
| B | 2, 2, 2 | 6 |
Comparing complete reward sequences by their totals.
What do you think happens?
Which outcome has the greater cumulative reward?
Reveal answer
Answer: Outcome B: 2, 2, 2
Outcome A totals 5, while Outcome B totals 6. Outcome B has the greater cumulative reward even though its first reward is smaller.
Mistakes in Reward Comparisons
Judging an outcome only by its immediate reward.
The reinforcement learning objective concerns cumulative reward over time, not only the next reward.
Fix:
Compare the complete reward sequences and determine their totals.Treating reward as a verbal description of the agent's purpose.
The reward signal is represented as a scalar number passed from the environment to the agent at each time step.
Fix:
Read each reward as a numerical signal and use the sequence of signals to measure the outcome.Stopping after inspecting one time step.
Rewards are accumulated across time before outcomes are compared.
Fix:
Include every reward in the sequence being evaluated.
When comparing outcomes, write each reward sequence in order, add the values across its time steps, and compare the resulting cumulative rewards. This keeps the immediate signal separate from the longer-run objective.
Practice with Reward Sequences
Outcome A produces rewards of 1, 3, and 1. Outcome B produces rewards of 2, 1, and 3. Which outcome has the greater cumulative reward?
Hints
- Add all three rewards for Outcome A.
- Add all three rewards for Outcome B.
- Compare the two totals rather than comparing only their first rewards.
Practice Answer
Compare Outcome A, with rewards 1, 3, and 1, with Outcome B, with rewards 2, 1, and 3.
Total Outcome A: The rewards 1, 3, and 1 give a cumulative reward of 5.
Total Outcome B: The rewards 2, 1, and 3 give a cumulative reward of 6.
Compare the totals: The total for Outcome B is greater than the total for Outcome A.
Outcome B has the greater cumulative reward.
Key Takeaways
- A reward is a scalar number passed from the environment to the agent at each time step.
- An immediate reward describes one time step, while cumulative reward considers rewards across multiple time steps.
- The reinforcement learning objective concerns cumulative reward over time rather than only the next reward.
- The reward hypothesis expresses an agent's purpose as maximizing the expected value of cumulative reward.
- To compare outcomes, evaluate their complete reward sequences and compare their totals.
Key Takeaways
- A reward is a scalar number sent from the environment to the agent at each time step.
- Immediate reward concerns the current time step; cumulative reward combines rewards over time.
- The reward hypothesis represents the agent's goal as maximizing the expected value of cumulative reward.
- The outcome with the greater cumulative reward is found by comparing complete reward sequences and their totals.