Concepts / Returns in Reinforcement Learning

Returns in Reinforcement Learning

Continuing tasks describe agent-environment interactions that do not naturally divide into identifiable episodes.

  • Programming

Why One Reward Is Not Enough

A reinforcement-learning agent receives rewards over time rather than usually receiving one single score. To describe what the agent is trying to optimize, we combine those rewards into a longer-term quantity called the return. The agent's objective is to maximize expected return, not merely one isolated reward.

A return is one quantity produced from a sequence of rewards by a specified function.

Tracing Rewards After Time t

Choose a time step called t. The rewards received after that point form an ordered sequence. The first reward after t is R t+1, followed by R t+2, R t+3, and so on. The ordering matters because each reward belongs to a particular point after t.

afternextnextcontinuestchosen time stepR t+1first reward after tR t+2second reward after tR t+3third reward after tR t+4 and latercontinuing sequence
Which rewards belong to the return sequence after time step t, and how are their positions indexed?

The return G t is defined from this ordered reward sequence by a specified function. It is not based on an unlabelled collection of rewards: the sequence begins after the selected time step and preserves the position of each reward.

Adding a Finite Reward Sequence

The simplest return function adds the rewards in the sequence. For one particular sequence, this produces one actual return.

A Three-Reward Sequence

A particular reward sequence is 2, 5, 1. Find the return when the return is defined as the sum of the rewards.

Identify the sequence: The rewards in order are 2, then 5, then 1.

Apply the sum definition: Add the rewards in the sequence: 2 + 5 + 1.

Evaluate the total: The total is 8.

The actual return for this reward sequence is 8.

applyproduces2, 5, 1ordered rewardsSumspecified function8actual return
How are the rewards in a sequence transformed into a single return value?

When Interaction Has No Endpoint

Some reinforcement-learning problems do not have a natural point at which one interaction ends and another begins. The agent and environment continue interacting without limit. These interactions are called continuing tasks.

This differs from an interaction that breaks naturally into identifiable episodes. In an episodic interaction, an episode provides a recognizable unit that ends before another interaction begins. A continuing task has no such natural division and instead keeps going.

divides intocontinues asEpisodicinteractionidentifiable episodesEpisodeends naturallyContinuing taskno natural endpointOngoing interactioncontinues without limit
How does an interaction that continues without a natural endpoint differ from one that breaks into identifiable episodes?

The Infinite-Horizon Problem

Because a continuing task has no final interaction step, its final time step is represented as T = ∞. The standard return formulation can therefore produce an infinite return.

What do you think happens?

Suppose the agent receives +1 at every time step in a continuing task. What happens when the return is the sum of all rewards after time t?

  • The return is 1
  • The return is a finite number
  • The return is infinite
Reveal answer

Answer: The return is infinite.

The interaction continues without limit, and every time step contributes +1. Adding the repeated rewards without an endpoint produces an infinite return.

nextnextcontinuessum produces+1t+1+1t+2+1t+3+1and every later step∞sum of rewards
What happens to the return when a task continues indefinitely and the agent receives +1 at every time step?

Actual Return and Expected Return

For one particular realized reward sequence, applying the return definition produces one actual return. The expected return is different: it is the expected value of the return. Reinforcement learning focuses on the expected quantity because the agent seeks to maximize expected return rather than merely one isolated outcome.

return function givesexpected value givesReward sequenceone particular sequenceActual returnone resulting valuePossible returnsreturns across outcomesExpected returnquantity to maximize
What is the difference between the return from one realized reward sequence and the expected return an agent aims to maximize?
QuantityWhat it describesRole
Actual returnThe result of applying the return definition to one particular reward sequenceDescribes one realized outcome
Expected returnThe expected value of the returnThe quantity the agent seeks to maximize

The two quantities answer different questions.

Common Reasoning Errors

  • Treating the return as only the next reward

    The return is defined from the sequence of rewards after t, not necessarily from only the first reward.

    Fix: First identify R t+1, R t+2, R t+3, and the remaining rewards included by the chosen return definition.

  • Ignoring reward order

    Each reward belongs to a particular point after t, so the sequence and its indexing matter.

    Fix: Write the rewards in order, beginning with R t+1.

  • Assuming every interaction has a final time step

    Continuing tasks do not naturally divide into identifiable episodes and continue without limit, with T represented as ∞.

    Fix: Check whether the interaction has a natural endpoint before treating its reward sequence as finite.

  • Confusing an actual return with expected return

    The value 8 is the actual return for one particular sequence. Expected return is the expected value of the return.

    Fix: Use actual return for one realized sequence and expected return for the quantity the agent seeks to maximize.

Practice the Distinction

EASY

A reward sequence after time t is 4, 0, and 3, and the return is defined as the sum of the sequence. State the actual return. Then explain whether that single value is the actual return or the expected return.

Hints
  • Add every reward in the ordered sequence.
  • Ask whether the question gives one realized sequence or information about possible outcomes.
MEDIUM

Decide whether each description is episodic or continuing: an interaction with a recognizable natural end, and an interaction that proceeds without a natural point where one interaction ends and another begins. Explain your choices.

Hints
  • Look for an identifiable episode boundary.
  • A continuing task has no natural endpoint and continues without limit.

Key Takeaways

  1. A return combines rewards received over time into one quantity using a specified function.
  2. After time t, the reward sequence begins with R t+1, R t+2, R t+3, and continues in order.
  3. The simplest return function adds the rewards in the sequence.
  4. Continuing tasks do not naturally divide into identifiable episodes and have no finite final time step; T is represented as ∞.
  5. One realized reward sequence produces an actual return, while the agent seeks to maximize expected return.

Key Takeaways

  • A return is a function that transforms an ordered sequence of rewards into one quantity.
  • The rewards after time t are R t+1, R t+2, R t+3, and so on.
  • Under the simplest return definition, the rewards are added together.
  • Continuing tasks have no natural endpoint, so repeated positive rewards can make the return infinite.
  • The actual return describes one realized sequence; expected return is the quantity the agent seeks to maximize.