Reward Signal
The reinforcement signal is the number that most directly guides changes in a policy or value function.
A Number That Changes Learning
In reinforcement learning, an agent does not change its policy or value function simply because a situation occurred. A numerical signal guides the size and direction of that change. This number is the reinforcement signal at time t. It may be positive, negative, or zero, and it is the number that most directly guides changes in a policy or value function.
The reinforcement signal is the major factor directing a learning update in response to the current situation. It acts as a multiplicative factor in parameter updates.
The TD Error Equation
The simplest case is an algorithm in which the reinforcement signal equals the reward signal Rₜ. In that case, the reward is the critical multiplier in the parameter-update equation. Other algorithms use a signal that contains more than the immediate reward. A TD error is one important example because it combines the reward with a change in predicted value.
δₜ = Rₜ + γV(Sₜ) − V(Sₜ₋₁)
Primary and Secondary Reinforcement
The reward term Rₜ is primary reinforcement because it is the direct reward contribution. The expression γV(Sₜ) − V(Sₜ₋₁) is secondary reinforcement because it represents the value-difference contribution. The distinction matters because a TD error can be present even when the reward term is zero: the value-difference contribution can still provide the secondary part of the signal.
Classifying a TD Error
Consider a TD error in which Rₜ is zero but γV(Sₜ) − V(Sₜ₋₁) is nonzero.
Inspect the reward term: Rₜ equals zero, so there is no primary reinforcement contribution.
Inspect the value-difference term: The value-difference term is nonzero, so the TD error still contains a secondary reinforcement contribution.
Classify the signal: Only the secondary contribution remains in the TD error.
The TD error represents pure secondary reinforcement.
The same expression can represent pure primary reinforcement, pure secondary reinforcement, or a mixture. These are not three different kinds of TD error. They are three ways to analyze which terms are contributing in a particular case.
| Case | Primary contribution | Secondary contribution | Interpretation |
|---|---|---|---|
| Pure primary | Rₜ is nonzero | γV(Sₜ) − V(Sₜ₋₁) is zero | The signal comes from the reward term alone |
| Pure secondary | Rₜ is zero | γV(Sₜ) − V(Sₜ₋₁) is nonzero | The signal comes from the value-difference term alone |
| Mixture | Rₜ contributes | γV(Sₜ) − V(Sₜ₋₁) contributes | Both parts are present |
Classifying the contents of a TD-error reinforcement signal
Reward as a Formal Goal
Reward signals are the reinforcement learning mechanism for formalizing goals. Because agents learn to maximize reward, the reward design determines what behavior is encouraged. A reward signal should communicate what to achieve rather than prescribe how to achieve it.
A chess-playing agent might be rewarded for taking opposing pieces or controlling the center of the board. Those actions can look useful, but they are not the same as winning the game. If the agent learns to maximize those rewards, it might pursue them even when doing so causes it to lose.
When evaluating a reward signal, ask whether maximizing it would reliably represent the outcome you actually want. The signal should reflect the actual desired outcome, not merely a convenient subgoal or a preferred method. It should communicate the goal itself and allow the agent to learn how to pursue it.
Mistakes in Reward Design
Treating a situation as if it automatically changes the policy or value function.
A numerical reinforcement signal guides the size and direction of the change.
Fix:
Identify the reinforcement signal that directs the update.Ignoring the secondary contribution when the immediate reward is zero.
The value-difference contribution can still be nonzero.
Fix:
Check both Rₜ and γV(Sₜ) − V(Sₜ₋₁).Rewarding a method or subgoal instead of the actual goal.
The agent learns to maximize the reward it receives, even if that behavior can lead away from the desired outcome.
Fix:
Design the signal to communicate what should be achieved, without prescribing how to achieve it.Treating primary and secondary reinforcement as separate learning signals.
They are conceptual parts of one TD-error reinforcement signal.
Fix:
Analyze the terms separately, then interpret their combination as δₜ.
Check Your Reasoning
A TD error contains a nonzero reward term and a zero value-difference term. Is the signal pure primary reinforcement, pure secondary reinforcement, or a mixture? Explain which contribution remains.
Hints
- Inspect Rₜ first.
- Then inspect γV(Sₜ) − V(Sₜ₋₁).
- Classify the signal according to which terms contribute.
Suppose a designer rewards an agent for performing an action that often helps achieve a goal, but does not reward the goal itself. What behavior might the agent learn to maximize, and what question should the designer ask about the reward signal?
Hints
- Agents learn to maximize the reward they receive.
- Compare the rewarded action with the actual desired outcome.
- Ask whether maximizing the reward reliably represents success.
The Essential Picture
- A reinforcement signal is the number that most directly guides changes in a policy or value function.
- In a TD error, Rₜ is the primary reinforcement contribution.
- The value-difference term γV(Sₜ) − V(Sₜ₋₁) is the secondary reinforcement contribution.
- A TD error can contain pure primary reinforcement, pure secondary reinforcement, or a mixture of both.
- Because agents learn to maximize received reward, reward design should represent the actual desired outcome rather than a convenient subgoal or preferred method.
Key Takeaways
- The reinforcement signal directs the size and direction of learning updates.
- A TD error combines immediate reward with a temporal-difference value contribution.
- Primary reinforcement is Rₜ; secondary reinforcement is γV(Sₜ) − V(Sₜ₋₁).
- The same TD-error expression can represent pure primary reinforcement, pure secondary reinforcement, or a mixture.
- A reward signal should formalize the actual goal, because the agent learns to maximize the reward it receives.