Cumulative Reward
Return turns a sequence of rewards into one quantity using a specified function.
Why One Reward Is Not Enough
A reinforcement-learning agent receives rewards over time rather than usually receiving one single score. To describe what the agent is trying to optimize, we combine those rewards into a longer-term quantity called the return. The agent seeks to maximize expected return, not merely one isolated reward.
Return turns a sequence of rewards into one quantity by applying a specified function to that sequence.
Choosing the Starting Time
Begin by choosing a time step called t. The rewards received after that point form an ordered sequence: R t+1, R t+2, R t+3, and so on. The first reward in the sequence is the reward at t+1, not the reward at t. The return G t is defined from this sequence by a specific function.
From Sequence to Return
Once the rewards after t have been identified, a specified function is applied to that ordered sequence. The result is one quantity: the return G t. There is more than one possible way to define this function. In the simplest case, the function adds the rewards together.
Return: the quantity produced by applying a specified function to the sequence of rewards received after a chosen time step.
Adding the Rewards
A Three-Reward Sequence
Use the simplest return definition for the reward sequence 2, 5, 1.
Identify the sequence: The three rewards are the rewards being considered after the chosen time step.
Start with the first reward: The running total begins at 2.
Add the second reward: Adding 5 to the running total gives 7.
Add the third reward: Adding 1 to 7 gives 8.
Under the sum definition, the actual return for the sequence 2, 5, 1 is 8.
With the simplest return definition, the order identifies the reward sequence, while addition combines that sequence into its total.
Actual and Expected Return
| Quantity | Meaning | Role |
|---|---|---|
| Actual return | The return produced from one particular observed reward sequence | Describes what that sequence produced |
| Expected return | The expected value of the return | The quantity the agent seeks to maximize |
For one particular reward sequence, applying the return definition produces one actual return. In the sequence 2, 5, 1, the actual return under the sum definition is 8. Expected return is different in emphasis: it is the expected value of the return, and this expected quantity is what the agent seeks to maximize.
Common Reasoning Errors
Including the reward at time step t in the sequence that starts after t.
The rewards after t are written as R t+1, R t+2, R t+3, and so on.
Fix:
Begin the sequence with R t+1.Treating the return as one isolated reward.
Return is defined from a sequence of rewards, not merely one isolated reward.
Fix:
First identify the ordered sequence after t, then apply the specified function.Assuming that every return must be calculated by addition.
A return is based on a specified function, and more than one way to define that function is possible.
Fix:
Use addition for the simplest case described here, while remembering that the function must be specified.Confusing an actual return with expected return.
The agent seeks to maximize expected return, not merely one isolated actual return.
Fix:
Use actual return for one particular sequence and expected return for the expected value of return.
Check Your Understanding
A chosen time step is t, and the rewards after that point are 3, 4, and 2. Under the simplest return definition, identify the reward sequence and calculate the actual return. Then state whether that result is an actual return or an expected return.
Hints
- The sequence begins with the first reward after t.
- For the simplest definition, add the rewards together.
- A result from one particular sequence is an actual return.
What do you think happens?
For the reward sequence 3, 4, 2, what is the actual return under the sum definition?
Reveal answer
Answer: 9
The simplest return is the sum of the rewards in the sequence: 3 + 4 + 2 = 9. Because this comes from one particular sequence, it is an actual return.
Key Takeaways
- Return turns an ordered sequence of rewards into one quantity using a specified function.
- When starting at time step t, the sequence begins with R t+1, followed by R t+2, R t+3, and so on.
- The simplest return is the sum of the rewards in the sequence.
- An actual return comes from one particular reward sequence.
- The agent seeks to maximize expected return.
Key Takeaways
- Return combines a sequence of rewards into one quantity through a specified function.
- The rewards after time step t are R t+1, R t+2, R t+3, and so on.
- Under the simplest definition, add the rewards to obtain the return.
- A particular sequence produces an actual return, while the agent seeks to maximize expected return.