Concepts / Cumulative Reward

Cumulative Reward

Return turns a sequence of rewards into one quantity using a specified function.

  • Programming

Why One Reward Is Not Enough

A reinforcement-learning agent receives rewards over time rather than usually receiving one single score. To describe what the agent is trying to optimize, we combine those rewards into a longer-term quantity called the return. The agent seeks to maximize expected return, not merely one isolated reward.

Return turns a sequence of rewards into one quantity by applying a specified function to that sequence.

Choosing the Starting Time

Begin by choosing a time step called t. The rewards received after that point form an ordered sequence: R t+1, R t+2, R t+3, and so on. The first reward in the sequence is the reward at t+1, not the reward at t. The return G t is defined from this sequence by a specific function.

afterordered nextordered nextcontinuestchosen time stepR t+1first reward after tR t+2next rewardR t+3next rewardlater rewardsand so on
Which rewards belong to the return when starting at time step t, and where does the sequence begin and end?

From Sequence to Return

Once the rewards after t have been identified, a specified function is applied to that ordered sequence. The result is one quantity: the return G t. There is more than one possible way to define this function. In the simplest case, the function adds the rewards together.

applyproduceReward sequenceR t+1, R t+2, R t+3, and soonSpecified functionfor example, additionReturn G tone quantity
How does a sequence of individual rewards become one return through a specified function?

Return: the quantity produced by applying a specified function to the sequence of rewards received after a chosen time step.

Adding the Rewards

A Three-Reward Sequence

Use the simplest return definition for the reward sequence 2, 5, 1.

Identify the sequence: The three rewards are the rewards being considered after the chosen time step.

Start with the first reward: The running total begins at 2.

Add the second reward: Adding 5 to the running total gives 7.

Add the third reward: Adding 1 to 7 gives 8.

Under the sum definition, the actual return for the sequence 2, 5, 1 is 8.

add 5add 12running total72 + 582 + 5 + 1
How do the rewards accumulate step by step to produce the total return?

With the simplest return definition, the order identifies the reward sequence, while addition combines that sequence into its total.

Actual and Expected Return

QuantityMeaningRole
Actual returnThe return produced from one particular observed reward sequenceDescribes what that sequence produced
Expected returnThe expected value of the returnThe quantity the agent seeks to maximize

For one particular reward sequence, applying the return definition produces one actual return. In the sequence 2, 5, 1, the actual return under the sum definition is 8. Expected return is different in emphasis: it is the expected value of the return, and this expected quantity is what the agent seeks to maximize.

apply return definitiontake expected valueOne reward sequence2, 5, 1Possible sequencesreturns considered togetherActual return8Expected returnagent seeks to maximize
What is the difference between the return from one observed reward sequence and the expected return the agent seeks to maximize?

Common Reasoning Errors

  • Including the reward at time step t in the sequence that starts after t.

    The rewards after t are written as R t+1, R t+2, R t+3, and so on.

    Fix: Begin the sequence with R t+1.

  • Treating the return as one isolated reward.

    Return is defined from a sequence of rewards, not merely one isolated reward.

    Fix: First identify the ordered sequence after t, then apply the specified function.

  • Assuming that every return must be calculated by addition.

    A return is based on a specified function, and more than one way to define that function is possible.

    Fix: Use addition for the simplest case described here, while remembering that the function must be specified.

  • Confusing an actual return with expected return.

    The agent seeks to maximize expected return, not merely one isolated actual return.

    Fix: Use actual return for one particular sequence and expected return for the expected value of return.

Check Your Understanding

EASY

A chosen time step is t, and the rewards after that point are 3, 4, and 2. Under the simplest return definition, identify the reward sequence and calculate the actual return. Then state whether that result is an actual return or an expected return.

Hints
  • The sequence begins with the first reward after t.
  • For the simplest definition, add the rewards together.
  • A result from one particular sequence is an actual return.

What do you think happens?

For the reward sequence 3, 4, 2, what is the actual return under the sum definition?

  • 6
  • 7
  • 9
  • 12
Reveal answer

Answer: 9

The simplest return is the sum of the rewards in the sequence: 3 + 4 + 2 = 9. Because this comes from one particular sequence, it is an actual return.

Key Takeaways

  1. Return turns an ordered sequence of rewards into one quantity using a specified function.
  2. When starting at time step t, the sequence begins with R t+1, followed by R t+2, R t+3, and so on.
  3. The simplest return is the sum of the rewards in the sequence.
  4. An actual return comes from one particular reward sequence.
  5. The agent seeks to maximize expected return.

Key Takeaways

  • Return combines a sequence of rewards into one quantity through a specified function.
  • The rewards after time step t are R t+1, R t+2, R t+3, and so on.
  • Under the simplest definition, add the rewards to obtain the return.
  • A particular sequence produces an actual return, while the agent seeks to maximize expected return.