Concepts / Secondary Reward

Secondary Reward

A reward signal is a numerical evaluation, not necessarily a physical reward.

  • Programming

From Event to Signal

Suppose an animal works to obtain food. In ordinary language, the food may be called a reward. In reinforcement learning, however, the learning system does not directly manipulate the food as its mathematical input. It receives a numerical evaluation, such as R_t, that scores what happened from the agent's learning perspective. That number may be positive, negative, or zero.

A reward is an object or event that an animal will approach and work for. A reward signal is a numerical evaluation used by a learning system to score an outcome.

approached and obtainedscores outcomeFoodR_tApproach and workLearning update
What is the difference between a physical event such as food and the numerical reward signal assigned to that event?

The physical event and the numerical signal can be related, but they are not the same thing. The signal is an abstraction used by the learning system.

Behavioral Meaning of Reward

In the context of animal behavior, an object or event counts as a reward when the animal tends to approach it and work to obtain it. This definition focuses on observable behavior rather than on whether people describe the object as pleasant, valuable, or a prize. The classification depends on what the animal does.

Food, sexual contact, and successful escape are examples of primary-reward-related outcomes in the source material. They are connected with survival or reproduction across an animal's ancestral history.

Primary and Secondary Origins

Primary and secondary rewards describe two different ways something can become rewarding. A primary reward comes from machinery built into an animal's nervous system through evolution. Primary rewards are therefore called innate. They are connected with outcomes that supported survival, reproduction, or predicted reproductive success across ancestral history.

A secondary reward acquires its rewarding quality during an individual's learning. The animal learns that a stimulus or event predicts a primary reward. A stimulus can also become secondary by predicting another learned reward. Its value is not supplied with the same built-in status as a primary reward; it is acquired through prediction.

shapes machinerysupportsacquirespredictscan supportEvolutionPrimary rewardinnateValued outcomeIndividual learningSecondary rewardpredicted value
How do innate primary rewards and learned secondary rewards connect within evolution and individual learning?
CategoryHow rewarding quality arisesSource example
Primary rewardEvolved nervous-system machineryTaste of nourishing food
Primary rewardConnected with ancestral survival or reproductionSexual contact
Primary rewardConnected with a beneficial outcomeSuccessful escape
Secondary rewardIndividual learning through predictionMoney

The distinction is based on origin: evolved machinery versus learned predictive value.

Acquiring Secondary Value

A previously neutral stimulus can become a secondary reward when an individual learns that it predicts a primary reward. The important change is not necessarily a change in the stimulus itself. The change is in what the stimulus predicts for that individual.

encountered withfollows or is predicted bylearned valueNeutral stimulusno learned predictionRepeated predictionSecondary rewardPrimary reward
How does a previously neutral stimulus become rewarding through repeated association with a primary reward?

Classifying Money

Why is money classified as a secondary reward rather than a primary reward?

Identify the category: Money is a secondary reward.

Identify the origin: Its rewarding quality is learned rather than supplied as innate primary-reward machinery.

Identify the prediction: Money acquires value through its predictive relationship with primary rewards or with other learned rewards.

Money is secondary because its rewarding quality is acquired through individual learning and prediction.

A complete classification includes both the category and the reason. Saying only that money is secondary is incomplete; the explanation must mention its learned predictive relationship with primary or other secondary rewards.

Reinforcement Directs Learning

A reinforcement signal is defined by its instructional role. It is the quantity that most directly directs changes in the agent's policy or value function in response to the current situation. In parameterized systems, it multiplicatively modulates parameter updates.

selectsproduces evaluationupdatesupdatesCurrent situationActionReinforcement signalPolicyValue function
How does a reward or TD error flow through a learning algorithm to change the agent's policy or value estimate?

The reward signal and reinforcement signal can coincide. If R_t is the critical multiplier in the parameter-update equation, then R_t is the reinforcement signal at time t. In other algorithms, the update-driving quantity contains more than the immediate reward.

Reading the TD Error

TD error = R_t + [γV(S_t) − V(S_{t−1})]

The TD error is therefore not necessarily just the reward received at the current moment. It can include an immediate reward contribution and a contribution from how predicted value changes. Because both terms can appear together, the TD error often represents a mixture of reinforcement sources.

combined withsubtracted fromaddedaddedR_timmediate rewardValue differenceTD errorreinforcement signalγV(S_t)discounted predicted valueV(S_{t−1})previous predicted value
How do the immediate reward and the difference between predicted and received future value combine to form the TD error?

Separating the Two TD Contributions

An update-driving signal contains an immediate reward and a change in predicted value. How should you interpret the two parts?

Read the immediate term: R_t reports the immediate reward contribution at time t.

Read the prediction term: γV(S_t) − V(S_{t−1}) reports how the predicted value changes between the relevant states.

Combine the terms: The TD error combines both contributions and can serve as the reinforcement signal that drives an update.

The TD error is a mixture of immediate reward and prediction-based value change, not necessarily an immediate reward alone.

A Careful Neuroscience Interpretation

The comparison between reinforcement signals and internal influences on decision making is useful, but it should not be taken too literally. A reward signal may be associated with something perceived in the outside world, yet internal experiences such as memories, ideas, or hallucinations can also trigger such a signal.

The mathematical R_t should therefore be treated as an abstraction for the combined effects of multiple neural signals. It is not proof that the brain contains one single signal identical to the mathematical quantity R_t.

Common Classification Errors

  • Treating the reward signal as the physical reward itself.

    R_t is a numerical evaluation, while the physical object or event is what the animal may approach and work to obtain.

    Fix: Describe the physical object or event separately from the numerical signal assigned to it.

  • Defining a reward only as something pleasant.

    The behavioral definition asks whether the animal approaches the object or event and works to obtain it.

    Fix: Use the animal's approach and effort as the relevant behavioral evidence.

  • Calling every positive number a reinforcement signal.

    The defining property of a reinforcement signal is its instructional role in directing updates, not whether its sign is positive, negative, or zero.

    Fix: Ask which quantity most directly modulates the learning update.

  • Calling money primary because people value it.

    The source identifies money as secondary because its rewarding quality is learned through prediction.

    Fix: Classify money as secondary and state that its value is acquired through a predictive relationship with primary or other secondary rewards.

  • Treating the TD error as only the immediate reward.

    The TD error combines R_t with γV(S_t) − V(S_{t−1}).

    Fix: Interpret the TD error as a mixture of immediate and prediction-based contributions.

Practice Classification

MEDIUM

For each case, identify whether the value is primary or secondary, and explain the origin of that value: the taste of nourishing food, money, a stimulus learned to predict money, and the numerical value R_t supplied to a learning algorithm.

Hints
  • Ask whether the value comes from evolved machinery or individual learning.
  • For R_t, ask whether it is an object or event or a numerical evaluation.
  • Give both the category and the reason.

What do you think happens?

Which statement best describes the TD error?

  • It is always identical to the physical reward object.
  • It contains an immediate reward contribution and a prediction-based value-change contribution.
  • It is any positive number produced by an agent.
  • It describes only whether an animal approaches an object.
Reveal answer

Answer: It contains an immediate reward contribution and a prediction-based value-change contribution.

The TD error combines R_t with γV(S_t) − V(S_{t−1}), so it can serve as a reinforcement signal containing both immediate and prediction-based contributions.

Key Takeaways

  1. A reward is an object or event that an animal approaches and works to obtain; a reward signal is a numerical evaluation.
  2. Primary rewards are innate because their rewarding quality comes from nervous-system machinery shaped by evolution.
  3. Secondary rewards acquire value through individual learning because they predict primary rewards or other learned rewards.
  4. A reinforcement signal is the quantity that most directly drives changes to a policy or value function.
  5. The TD error combines immediate reward with a change in predicted value and can therefore act as a mixed reinforcement signal.

Key Takeaways

  • A physical reward and a reward signal are different: the first is an object or event, while the second is a numerical evaluation.
  • Primary reward is innate and reflects evolved mechanisms connected with survival or reproduction.
  • Secondary reward is learned when a stimulus or event predicts a primary or another learned reward.
  • The reinforcement signal directly guides policy or value-function updates.
  • The TD error combines immediate reward with prediction-based value change.