Concepts / Understanding Reward Signals

Understanding Reward Signals

A reward is a scalar number passed from the environment to the agent at each time step.

  • Programming

Why One Number Matters

In reinforcement learning, an agent can be described informally as having a purpose: it should achieve some desirable outcome. A reward signal turns that purpose into something measurable. At each time step, the environment passes the agent a scalar number called a reward. The agent is then understood to seek more reward over the long run, rather than judging success from only one number at one moment.

suppliespasses at each time stepEnvironmentRewardscalar numberAgent
How does a scalar reward move from the environment to the agent at each time step?

A reward is not a paragraph describing what the agent should do. It is a scalar number supplied by the environment at each time step.

From One Reward to a Sequence

The reward received at the current time step is an immediate reward. It tells the agent about one point in the interaction. Reinforcement learning, however, is concerned with cumulative reward over time. That means the rewards from multiple time steps are considered together as a sequence, and the sequence is judged by its total rather than by only its first or next reward.

next time stepnext time stepaccumulateTime 1reward 2Time 2reward 3Time 3reward 1Cumulative reward6
How does a reward at the current time step differ from the total accumulated across several time steps?

Adding a Short Reward Sequence

An outcome produces rewards of 2, 3, and 1 at three successive time steps. What is its cumulative reward?

List the rewards: The sequence contains three scalar rewards: 2, 3, and 1.

Combine the time steps: Consider the rewards together instead of judging only the first or next reward.

Calculate the total: Adding 2, 3, and 1 gives a cumulative reward of 6.

The outcome has a cumulative reward of 6.

The Reward Hypothesis

The reward hypothesis formalizes the agent's purpose as maximizing the expected value of cumulative reward. This connects the three central ideas: the environment supplies a simple number at each time step, those numbers are accumulated across time, and the agent's goal is expressed through the expected value of that accumulated result.

becomes measurable throughis accumulated intodefinesPurposedesirable outcomeReward signalscalar numbersCumulative rewardacross timeAgent objectivemaximize expected value
How does the reward hypothesis turn an informal purpose into a measurable reinforcement learning objective?

Comparing Complete Outcomes

To apply the reward hypothesis, compare complete reward sequences and their totals. A sequence that begins with a larger immediate reward does not necessarily produce the larger cumulative reward. The relevant comparison is the total across the time steps being considered.

OutcomeReward sequenceCumulative reward
A4, 0, 15
B2, 2, 26

Comparing complete reward sequences by their totals.

totaltotalOutcome A4, 0, 1Outcome B2, 2, 25cumulative reward6cumulative reward
Given two short reward sequences, which outcome has the greater cumulative reward?

What do you think happens?

Which outcome has the greater cumulative reward?

  • Outcome A: 4, 0, 1
  • Outcome B: 2, 2, 2
  • They are equal
Reveal answer

Answer: Outcome B: 2, 2, 2

Outcome A totals 5, while Outcome B totals 6. Outcome B has the greater cumulative reward even though its first reward is smaller.

Mistakes in Reward Comparisons

  • Judging an outcome only by its immediate reward.

    The reinforcement learning objective concerns cumulative reward over time, not only the next reward.

    Fix: Compare the complete reward sequences and determine their totals.

  • Treating reward as a verbal description of the agent's purpose.

    The reward signal is represented as a scalar number passed from the environment to the agent at each time step.

    Fix: Read each reward as a numerical signal and use the sequence of signals to measure the outcome.

  • Stopping after inspecting one time step.

    Rewards are accumulated across time before outcomes are compared.

    Fix: Include every reward in the sequence being evaluated.

When comparing outcomes, write each reward sequence in order, add the values across its time steps, and compare the resulting cumulative rewards. This keeps the immediate signal separate from the longer-run objective.

Practice with Reward Sequences

EASY

Outcome A produces rewards of 1, 3, and 1. Outcome B produces rewards of 2, 1, and 3. Which outcome has the greater cumulative reward?

Hints
  • Add all three rewards for Outcome A.
  • Add all three rewards for Outcome B.
  • Compare the two totals rather than comparing only their first rewards.

Practice Answer

Compare Outcome A, with rewards 1, 3, and 1, with Outcome B, with rewards 2, 1, and 3.

Total Outcome A: The rewards 1, 3, and 1 give a cumulative reward of 5.

Total Outcome B: The rewards 2, 1, and 3 give a cumulative reward of 6.

Compare the totals: The total for Outcome B is greater than the total for Outcome A.

Outcome B has the greater cumulative reward.

Key Takeaways

  1. A reward is a scalar number passed from the environment to the agent at each time step.
  2. An immediate reward describes one time step, while cumulative reward considers rewards across multiple time steps.
  3. The reinforcement learning objective concerns cumulative reward over time rather than only the next reward.
  4. The reward hypothesis expresses an agent's purpose as maximizing the expected value of cumulative reward.
  5. To compare outcomes, evaluate their complete reward sequences and compare their totals.

Key Takeaways

  • A reward is a scalar number sent from the environment to the agent at each time step.
  • Immediate reward concerns the current time step; cumulative reward combines rewards over time.
  • The reward hypothesis represents the agent's goal as maximizing the expected value of cumulative reward.
  • The outcome with the greater cumulative reward is found by comparing complete reward sequences and their totals.