Rewards in Reinforcement Learning
Return turns a sequence of rewards into one quantity using a specified function.
From Rewards to a Goal
A reinforcement-learning agent receives rewards over time rather than usually receiving one single score. That creates a modeling question: how can a sequence of separate rewards describe what the agent is trying to optimize? Reinforcement learning uses the return for this purpose. The return turns a sequence of rewards into one quantity by applying a specified function to that sequence.
The agent seeks to maximize expected return, not merely one isolated reward.
Selecting Rewards After t
Choose a time step called t. The rewards received after that point form an ordered sequence: R t+1, R t+2, R t+3, and so on. These are the rewards that belong to the return beginning after time step t. The ordering matters because each reward belongs to a particular point after t. The return at time t is written as G t and is defined from this reward sequence by a specified function.
Turning a Sequence into One Return
A return is not the reward sequence itself. It is the result of applying a specified function to that sequence. More than one function could be used to define a return. In the simplest case, the function adds the rewards together, so the return is the total of the rewards in the sequence.
Summing Three Rewards
Using the simplest return definition, find the return for the reward sequence 2, 5, 1.
Identify the sequence: The rewards in this particular sequence are 2, 5, and 1.
Apply the sum definition: Add the rewards in the sequence: 2 + 5 + 1.
Compute the result: The total is 8.
The actual return for this particular sequence is 8.
Actual and Expected Return
For one particular reward sequence, applying the return definition produces one actual return. In the worked example, the sequence 2, 5, 1 produces an actual return of 8 under the sum definition. Expected return is a different quantity: it is the expected value of the return. The agent's objective is to maximize this expected quantity rather than to maximize one isolated reward or to focus only on one already-observed sequence.
Reward Signals Define the Problem
The reward signal helps define the reinforcement-learning problem together with the environment. It provides the rewards that are received over time, and those rewards are combined into a return so that the agent's longer-term objective can be described. Without confusing the reward sequence with the return, the relationship is: the environment and the problem provide a reward signal, the signal produces a sequence of rewards, and a specified return function turns that sequence into the quantity whose expected value the agent seeks to maximize.
What R_t Represents
R_t is the reward signal at time t. It is a theoretical abstraction used to describe the reinforcement-learning problem and the rewards available to the agent at a particular time.
The subscript t identifies the time at which the signal is being described. When discussing the return after time step t, the relevant later rewards are written R t+1, R t+2, R t+3, and so on. Thus, R_t names the signal at the selected time, while the return beginning after that point is built from the later reward sequence.
Reward and Reinforcement Signals
Reward signals and reinforcement signals are distinct concepts in reinforcement learning theory. The reward signal is the theoretical signal used to help define the problem and the rewards available over time. A reinforcement signal is a separate theoretical concept concerning the signal that influences learning or behavior. Keeping these concepts distinct prevents the word reward from being treated as though it names every signal involved in learning.
The Biological Interpretation
Neural signals are physiological events that may behave like theoretical signals in function. This means that an abstract signal in a model can be related to biological processes without being identical to one specific physical event.
Mistakes with Reward Sequences
Including R_t when selecting the rewards after time step t.
The rewards after time step t are written R t+1, R t+2, R t+3, and so on.
Fix:
Start the after-t sequence with R t+1.Treating the return as one isolated reward.
The return combines a sequence of rewards into one quantity.
Fix:
Identify the relevant sequence first, then apply the specified return function.Confusing an actual return with expected return.
The value 8 is the actual return for that particular sequence under the sum definition. Expected return is the expected value of return.
Fix:
Use actual return for one particular sequence and expected return for the quantity the agent seeks to maximize.Treating R_t as a literal unitary master reward signal in an animal's brain.
R_t is a theoretical reward signal, while biological systems may involve many neural signals and many systems.
Fix:
Describe R_t as a model-level abstraction that may relate functionally to physiological signals without being identical to one specific event.
Check Your Understanding
A chosen time step is t, and the later reward sequence is 4, 0, and 3. Under the simplest return definition, identify the rewards that belong to the sequence, calculate the actual return, and state whether that value is automatically the agent's expected return.
Hints
- The sequence after t begins with R t+1.
- Use the sum definition for the simplest return.
- Expected return is the expected value of return, not automatically the actual return from one sequence.
Practice Solution
For the later reward sequence 4, 0, and 3, use the simplest return definition.
Name the sequence: The rewards after t are R t+1, R t+2, and R t+3, with values 4, 0, and 3.
Sum the rewards: Add 4, 0, and 3.
Interpret the result: The actual return for this particular sequence is 7. It is not automatically the expected return.
The actual return is 7, while expected return refers to the expected value of return that the agent seeks to maximize.
Key Takeaways
- The return turns an ordered reward sequence into one quantity using a specified function.
- After time step t, the sequence begins with R t+1, followed by R t+2, R t+3, and so on.
- Under the simplest definition, the return is the sum of the rewards in the sequence.
- An actual return comes from one particular sequence, while expected return is the expected value of return that the agent seeks to maximize.
- R_t is a theoretical reward signal, not automatically one unitary physiological signal in an animal's brain.
Key Takeaways
- Return is a function that maps a sequence of rewards to one quantity.
- The rewards after time step t are R t+1, R t+2, R t+3, and so on.
- The simplest return is the sum of the rewards in the sequence.
- The agent seeks to maximize expected return rather than one isolated reward or one particular actual return.
- R_t is a theoretical reward signal and should not be treated as a single unchanged physiological master signal in an animal's brain.