Primary Reward
A reward signal is a numerical evaluation, not necessarily a physical reward.
The Number Behind a Reward
When people hear the word reward, they may imagine food, money, praise, or an event that an animal approaches. Reinforcement learning uses the term more precisely. At time t, an agent receives an evaluative number called R_t. This reward signal may be positive, negative, or zero, and it tells the learning system how valuable or undesirable what happened was from the task's perspective.
A reward signal is a numerical evaluation, not necessarily a physical reward.
Primary and Secondary Sources
Primary reward and secondary reward describe two ways something can become rewarding. Primary reward comes from mechanisms built into an animal's nervous system through evolution. Examples in the source include the taste of nourishing food, sexual contact, and successful escape. These outcomes are connected with survival and reproduction across an animal's ancestral history.
Secondary reward is acquired during an individual's learning. A stimulus or event becomes rewarding because it predicts primary reward or predicts another learned reward. Money is the central example: its rewarding quality is presented as something conferred through learning rather than as primary reward machinery.
Classifying Two Kinds of Reward
Use the source distinction to classify nourishing food and money.
Nourishing food: The source presents its taste as an example of primary reward connected to outcomes that supported survival across ancestral history.
Money: The source presents money as a secondary reward whose rewarding quality is acquired through learning because it predicts primary reward or another learned reward.
Primary reward is connected with evolved machinery; secondary reward is acquired through prediction and learning.
What Actually Changes
A reinforcement signal is identified by the job it performs, not simply by whether its value is positive, negative, or zero. It is the quantity that most directly directs changes to the policy or value function in response to the current situation. In more specific terms, it multiplicatively modulates parameter updates.
The reward signal and reinforcement signal can be the same quantity. If an algorithm uses R_t as the critical multiplier in its parameter-update equation, then R_t is the reinforcement signal at time t. In other algorithms, the update-driving quantity contains more than the immediate reward, so the two terms should not be treated as automatically identical.
Reading the TD Error
A temporal-difference error is an important example of a reinforcement signal that may contain more than the immediate reward. It combines the immediate reward R_t with a change in predicted value. The prediction-based contribution is written as gamma times V(S_t) minus V(S_{t-1}). Thus, the TD error can represent a mixture of a primary reinforcement contribution from the immediate reward and a secondary reinforcement contribution from the temporal difference in predicted values.
Separating the Two TD Contributions
Suppose an update-driving TD error contains an immediate reward contribution of 3 and a prediction-based contribution of negative 1. Interpret the two parts without treating the TD error as only an immediate reward.
Immediate contribution: The value 3 represents the R_t portion: the primary reinforcement contribution supplied by the immediate reward.
Prediction-based contribution: The value negative 1 represents the change in predicted value, the secondary reinforcement contribution described by gamma V(S_t) minus V(S_{t-1}).
Combined signal: Together, the contributions produce a TD error of 2. The update-driving quantity therefore contains both what happened immediately and how predicted value changed.
A TD error can be a mixture of immediate and prediction-based reinforcement contributions.
What do you think happens?
If an algorithm's update-driving quantity contains an immediate reward and a change in predicted value, is that quantity necessarily identical to the immediate reward?
Reveal answer
Answer: No, because the update-driving quantity can include the prediction-based contribution.
The source identifies the TD error as a reinforcement signal that combines R_t with the temporal difference in predicted values.
A Careful Neuroscience Analogy
The reward-signal idea can be compared with an internal influence on decision making and learning, but the comparison should not be taken literally. A signal may be associated with something perceived in the outside world, while internal experiences such as memories, ideas, or hallucinations can also trigger such a signal.
The mathematical R_t is best treated as an abstraction for the combined effects of multiple neural signals. It should not be interpreted as proof that the brain contains one single signal identical to the mathematical quantity R_t.
Common Interpretation Errors
Treating reward as only a physical object or event
The technical reward signal is a numerical evaluation, represented here by R_t.
Fix:
Separate the object or event from the numerical signal that evaluates what happened.Assuming every positive number is a reinforcement signal
A reinforcement signal is defined by its instructional role in directing updates, including its multiplicative modulation of parameter updates.
Fix:
Ask whether the quantity most directly drives the relevant update.Equating primary reward with anything rewarding
The source presents money as secondary reward acquired through learning because it predicts primary reward or another learned reward.
Fix:
Check whether the rewarding quality is connected to evolved machinery or acquired through prediction.Reducing the TD error to the immediate reward
The TD error can also contain the change in predicted value, written as gamma V(S_t) minus V(S_{t-1}).
Fix:
Identify both the immediate and prediction-based contributions.
Practice: Classify the Signal
For each description, identify the most precise concept: primary reward, secondary reward, reward signal, or reinforcement signal. First, a value is acquired because it predicts a primary reward. Second, a numerical evaluation R_t is delivered after an event. Third, a quantity directly modulates a parameter update. Fourth, a TD error combines an immediate reward with a change in predicted value.
Hints
- Look for the distinction between evolution and individual learning.
- A numerical evaluation is not the same thing as the object or event that may produce it.
- Focus on which quantity directly drives the update.
- Check whether the signal contains both R_t and a prediction-based contribution.
A reliable method is to ask two questions: What is the source of the rewarding quality, and what quantity directly drives the learning update?
Key Takeaways
- A reward signal is a numerical evaluation, not necessarily a physical reward object or event.
- Primary reward comes from evolved mechanisms, while secondary reward is acquired through learning and prediction.
- A reinforcement signal is identified by its direct role in modulating policy or value-function updates.
- The immediate reward and reinforcement signal can coincide, but they are not always identical.
- A TD error can combine the immediate reward R_t with a change in predicted value, making it a mixed reinforcement signal.
Key Takeaways
- Reward signals are numerical evaluations used by a learning system.
- Primary reward is connected with evolved mechanisms; secondary reward is acquired through individual learning and prediction.
- The reinforcement signal is the quantity that most directly directs policy or value-function updates.
- A TD error can contain both an immediate reward contribution and a prediction-based contribution.