Returns in Reinforcement Learning
Continuing tasks describe agent-environment interactions that do not naturally divide into identifiable episodes.
Why One Reward Is Not Enough
A reinforcement-learning agent receives rewards over time rather than usually receiving one single score. To describe what the agent is trying to optimize, we combine those rewards into a longer-term quantity called the return. The agent's objective is to maximize expected return, not merely one isolated reward.
A return is one quantity produced from a sequence of rewards by a specified function.
Tracing Rewards After Time t
Choose a time step called t. The rewards received after that point form an ordered sequence. The first reward after t is R t+1, followed by R t+2, R t+3, and so on. The ordering matters because each reward belongs to a particular point after t.
The return G t is defined from this ordered reward sequence by a specified function. It is not based on an unlabelled collection of rewards: the sequence begins after the selected time step and preserves the position of each reward.
Adding a Finite Reward Sequence
The simplest return function adds the rewards in the sequence. For one particular sequence, this produces one actual return.
A Three-Reward Sequence
A particular reward sequence is 2, 5, 1. Find the return when the return is defined as the sum of the rewards.
Identify the sequence: The rewards in order are 2, then 5, then 1.
Apply the sum definition: Add the rewards in the sequence: 2 + 5 + 1.
Evaluate the total: The total is 8.
The actual return for this reward sequence is 8.
When Interaction Has No Endpoint
Some reinforcement-learning problems do not have a natural point at which one interaction ends and another begins. The agent and environment continue interacting without limit. These interactions are called continuing tasks.
This differs from an interaction that breaks naturally into identifiable episodes. In an episodic interaction, an episode provides a recognizable unit that ends before another interaction begins. A continuing task has no such natural division and instead keeps going.
The Infinite-Horizon Problem
Because a continuing task has no final interaction step, its final time step is represented as T = ∞. The standard return formulation can therefore produce an infinite return.
What do you think happens?
Suppose the agent receives +1 at every time step in a continuing task. What happens when the return is the sum of all rewards after time t?
Reveal answer
Answer: The return is infinite.
The interaction continues without limit, and every time step contributes +1. Adding the repeated rewards without an endpoint produces an infinite return.
Actual Return and Expected Return
For one particular realized reward sequence, applying the return definition produces one actual return. The expected return is different: it is the expected value of the return. Reinforcement learning focuses on the expected quantity because the agent seeks to maximize expected return rather than merely one isolated outcome.
| Quantity | What it describes | Role |
|---|---|---|
| Actual return | The result of applying the return definition to one particular reward sequence | Describes one realized outcome |
| Expected return | The expected value of the return | The quantity the agent seeks to maximize |
The two quantities answer different questions.
Common Reasoning Errors
Treating the return as only the next reward
The return is defined from the sequence of rewards after t, not necessarily from only the first reward.
Fix:
First identify R t+1, R t+2, R t+3, and the remaining rewards included by the chosen return definition.Ignoring reward order
Each reward belongs to a particular point after t, so the sequence and its indexing matter.
Fix:
Write the rewards in order, beginning with R t+1.Assuming every interaction has a final time step
Continuing tasks do not naturally divide into identifiable episodes and continue without limit, with T represented as ∞.
Fix:
Check whether the interaction has a natural endpoint before treating its reward sequence as finite.Confusing an actual return with expected return
The value 8 is the actual return for one particular sequence. Expected return is the expected value of the return.
Fix:
Use actual return for one realized sequence and expected return for the quantity the agent seeks to maximize.
Practice the Distinction
A reward sequence after time t is 4, 0, and 3, and the return is defined as the sum of the sequence. State the actual return. Then explain whether that single value is the actual return or the expected return.
Hints
- Add every reward in the ordered sequence.
- Ask whether the question gives one realized sequence or information about possible outcomes.
Decide whether each description is episodic or continuing: an interaction with a recognizable natural end, and an interaction that proceeds without a natural point where one interaction ends and another begins. Explain your choices.
Hints
- Look for an identifiable episode boundary.
- A continuing task has no natural endpoint and continues without limit.
Key Takeaways
- A return combines rewards received over time into one quantity using a specified function.
- After time t, the reward sequence begins with R t+1, R t+2, R t+3, and continues in order.
- The simplest return function adds the rewards in the sequence.
- Continuing tasks do not naturally divide into identifiable episodes and have no finite final time step; T is represented as ∞.
- One realized reward sequence produces an actual return, while the agent seeks to maximize expected return.
Key Takeaways
- A return is a function that transforms an ordered sequence of rewards into one quantity.
- The rewards after time t are R t+1, R t+2, R t+3, and so on.
- Under the simplest return definition, the rewards are added together.
- Continuing tasks have no natural endpoint, so repeated positive rewards can make the return infinite.
- The actual return describes one realized sequence; expected return is the quantity the agent seeks to maximize.