Secondary Reinforcement
Eligibility traces help reinforcement learning algorithms address delayed reinforcement.
Why Delay Disrupts Learning
In instrumental conditioning, an action and its reward do not always occur together. When the reward arrives later, it becomes more difficult for the reward to influence learning about the earlier action or situation. The learning system must somehow connect the earlier event with the later reinforcing outcome.
An action followed by a delayed outcome
Trace what happens when an action is followed by a delay and then by a primary reward.
Action: An action occurs and becomes relevant to later learning.
Delay: Time passes before the primary reward arrives. The longer interval makes it harder for the later reward to influence learning about the earlier action.
Primary reward: The reinforcing outcome arrives, but the learning system must still address the separation between this outcome and the earlier action.
Delayed reinforcement creates a credit-assignment problem: the later outcome must influence an earlier event across the delay.
Eligibility Across the Delay
An eligibility trace is a mechanism within a reinforcement learning algorithm for addressing delayed reinforcement. It can be understood as a connection between an earlier learning-relevant event and reinforcement that arrives later. Rather than allowing the delay to erase the earlier event's relevance, the algorithm uses the trace to preserve which earlier states or actions should receive credit when reinforcement finally appears.
The trace does not mean that the earlier event and the later reward occur at the same time. It is an algorithmic mechanism for dealing with their separation. Eligibility traces therefore address the delay itself: they help preserve the earlier event's relevance until a later reinforcing outcome can affect learning.
From Stimulus Traces to Eligibility Traces
Early theories of animal learning used the idea of stimulus traces. Reinforcement learning uses eligibility traces. These are not presented as identical mechanisms or as interchangeable terms. Their relationship is a historical and functional comparison: both ideas help explain how learning can remain connected to an earlier event when an important consequence is separated from it in time.
Evaluative Feedback and Value
Eligibility traces are one mechanism for the delayed-reinforcement problem. The other mechanism identified in the source is the value function learned via temporal-difference algorithms. Value functions provide nearly immediate evaluative feedback about an ongoing situation or outcome. In this comparison, value functions correspond to the role of secondary reinforcement, while eligibility traces address how learning can reach backward across the delay.
| Mechanism or idea | Main role in the delay problem | Context |
|---|---|---|
| Eligibility trace | Connects an earlier learning-relevant event with later reinforcement | Reinforcement-learning algorithm |
| Value function | Provides nearly immediate evaluative feedback | Temporal-difference reinforcement learning |
| Secondary reinforcement | Provides a more immediate reinforcing influence during the delay | Instrumental conditioning |
| Stimulus trace | Helps explain lingering influence of a stimulus across time | Early theories of animal learning |
Secondary Reinforcement During Delay
Secondary reinforcement is a more immediate reinforcing influence that supports learning during the interval between an action and a delayed primary reward. In Hull's account, learning can be supported not only by the primary reward at the end, but also by secondary reinforcement that occurs during the delay.
Regularly occurring stimuli during the delay can favor the development of secondary reinforcement. Because these stimuli appear during the interval, they can provide a reinforcing influence before the primary reward arrives. The learner therefore does not have to rely only on the final reward to span the entire delay.
The important idea is not simply that a delay exists. The effect of the delay also depends on what happens during it. Repeated or regularly occurring stimuli can bridge part of the interval by supporting secondary reinforcement.
Extending the Goal Gradient
Molar stimulus traces refer to the lingering presence of a stimulus in Hull's theory. The source presents them as part of an explanation for how a goal gradient can span time. Hull's proposal is that secondary reinforcement works together with these traces, allowing the resulting gradient to extend beyond the period that stimulus traces alone would cover.
This account explains why secondary reinforcement matters for a distant goal. Stimulus traces provide a lingering influence, while secondary reinforcement supplies additional support during the delay. Together, they can allow the goal gradient to span more of the time separating action from reward.
When the Bridge Is Blocked
| Condition | What occurs during the delay | Effect of increasing delay |
|---|---|---|
| Secondary reinforcement supported | Regularly occurring stimuli can favor secondary reinforcement | Learning decreases with increased delay less |
| Secondary reinforcement obstructed | The conditions that support secondary reinforcement are obstructed | Learning decreases with increased delay more |
Animal experiments described in the source support this distinction: when conditions favor secondary reinforcement during a delay, learning decreases with increased delay less than it does when those conditions obstruct secondary reinforcement. Thus, a delayed reward is not the complete explanation of performance. The conditions inside the delay also matter.
Common Conceptual Mistakes
Treating eligibility traces and stimulus traces as identical.
The source says that eligibility traces are similar to stimulus traces, but places them in different theoretical contexts.
Fix:
Describe the relationship as a functional and historical similarity: both help explain learning across a time separation.Treating value functions and eligibility traces as the same mechanism.
The source assigns different roles to them.
Fix:
Associate eligibility traces with addressing the delay and value functions with nearly immediate evaluative feedback.Assuming that every delay has the same learning effect.
The source emphasizes that the effect depends partly on whether conditions support or obstruct secondary reinforcement.
Fix:
Compare delays while also examining what regularly occurs during the interval.Defining secondary reinforcement as the final primary reward.
Secondary reinforcement is the more immediate reinforcing influence during the delay, distinct from the primary reward at the end.
Fix:
Identify the regularly occurring stimuli or evaluative influences that support learning before the primary reward arrives.
Check Your Understanding
A delayed primary reward follows an action. In one condition, regularly occurring stimuli appear during the delay. In another condition, the opportunities for those stimuli to support secondary reinforcement are obstructed. Explain why the two conditions can produce different amounts of learning as the delay increases. Then identify which part of the explanation belongs to secondary reinforcement and which part belongs to eligibility traces.
Hints
- First compare what happens during the two delays.
- Then explain how a more immediate reinforcing influence can support learning before the primary reward.
- Finally distinguish the algorithmic role of an eligibility trace from the role of secondary reinforcement.
What do you think happens?
If the delay between an action and a primary reward increases, will learning necessarily decline by the same amount when secondary reinforcement is supported and when it is obstructed?
Reveal answer
Answer: No, learning decreases with increased delay less when secondary reinforcement is supported.
The source distinguishes delays that favor secondary reinforcement from delays that obstruct it. Regularly occurring stimuli can provide a more immediate reinforcing influence during the interval.
Key Takeaways
- Delayed reinforcement makes learning more difficult because an earlier action or situation must be connected with a later outcome.
- Eligibility traces are reinforcement-learning mechanisms that preserve the relevance of earlier events until later reinforcement arrives.
- Eligibility traces are similar to, but not identical with, stimulus traces from early theories of animal learning.
- Value functions learned through temporal-difference algorithms provide nearly immediate evaluative feedback and correspond to the role of secondary reinforcement.
- Regularly occurring stimuli during a delay can support secondary reinforcement, extend a goal gradient, and reduce the decline in learning caused by increased delay.
- Learning is affected differently when secondary reinforcement is supported than when the conditions for it are obstructed.
Key Takeaways
- Delayed reinforcement creates a temporal connection problem between an earlier event and a later outcome.
- Eligibility traces help reinforcement-learning algorithms preserve the relevance of earlier states or actions across that delay.
- Secondary reinforcement supplies a more immediate reinforcing influence during the interval before the primary reward.
- Stimulus traces and eligibility traces are functionally similar ideas from different theoretical contexts.
- The effect of delay depends on whether conditions support or obstruct secondary reinforcement.