Eligibility Traces in Reinforcement Learning
Miller's 1981 Law-of-Effect-like rule included synaptically local contingent eligibility traces.
The Delayed-Reward Problem
A reinforcement-learning system may need to connect two events separated in time: an earlier synaptic event and a later reinforcement signal. The central question is how the system can preserve enough information about the earlier event for the later reward to affect the appropriate synapse. Miller's 1981 Law-of-Effect-like proposal addressed this problem with synaptically local contingent eligibility traces.
An eligibility trace is the temporary synaptic record that makes an earlier event eligible to be changed when a later reinforcement signal arrives.
A Synapse Becomes Eligible
In Miller's proposal, eligibility is local to a synapse. An earlier sensory or synaptic event does not immediately need to produce the final learning change. Instead, that event can leave the synapse temporarily eligible. If a reinforcement signal arrives afterward, the signal can be related to the synapse that was previously marked. The trace therefore addresses temporal credit assignment: it helps connect a later outcome with an earlier event.
Tracing One Earlier Event
Suppose a sensory event is followed by activity at one synapse, and a reinforcement signal arrives later. How does a contingent eligibility trace organize the learning sequence?
Earlier event: The sensory event and associated synaptic activity occur first.
Temporary eligibility: The synapse retains a local eligibility trace rather than being treated as unrelated to the later outcome.
Later reinforcement: A reinforcement signal arrives after the earlier synaptic event.
Credit assignment: The later signal can affect the synapse because that synapse was previously made eligible.
The trace supplies a time-bridging link between the earlier synaptic event and the later reinforcement signal.
Miller's Sensory Analyzer
Miller's proposal also included a sensory analyzer unit with a critic-like role. The important function is not merely sensing input. The analyzer supplied reinforcement signals to neurons, providing the signal that could guide changes at synapses that had become eligible. This anticipates a central feature later associated with actor-critic architectures: one part of the system contributes behavior, while a critic-like mechanism supplies information related to reinforcement.
The Hedonistic Synapse
Seung's hedonistic synapse gives the learning problem a synapse-level interpretation. The synapse changes its neurotransmitter release probability according to which event is followed by reward: release or failure to release. The relevant distinction is between synaptic activity and the probability of the synapse's future neurotransmitter release. A hedonistic synapse therefore expresses learning as a reward-dependent change in release probability, rather than treating every synaptic event as an identical outcome.
Release Versus Failure to Release
What does the hedonistic-synapse idea track when reward follows a synaptic event?
Observe the synaptic event: The synapse either releases neurotransmitter or fails to release it.
Observe the later outcome: The system determines whether reward follows that release or failure to release.
Adjust the synapse: The synapse changes its neurotransmitter release probability according to the relationship between the event and the reward.
The learned variable is release probability, and its adjustment depends on whether reward followed release or failure to release.
Historical Connections
| Idea | Earlier event | Later signal | Learning consequence |
|---|---|---|---|
| Contingent eligibility trace | Earlier synaptic event | Reinforcement signal | The locally eligible synapse can be affected |
| Reward-modulated STDP | Synaptic activity | Reward-related modulation | The source connects it with the general features of eligibility-based reward learning |
| Actor-critic method | Past behavior or synaptic event | TD error | The source identifies TD error as a later actor-critic development |
The shared question across these ideas is how to connect a past event with a later reinforcement signal. The source connects Miller's proposal with the general features of reward-modulated STDP and connects the later actor-critic architecture with TD error. The useful comparison is therefore structural: an earlier synaptic or behavioral event is related to an eligibility-like record, a later reward-related signal arrives, and learning uses the relationship between the two. This relationship should not be mistaken for proof that all three approaches implement the same mechanism.
Bacterial chemotaxis belongs here as a metaphor for the problem of using consequences to guide future behavior. It should not be confused with the neural learning mechanisms themselves. The mechanisms discussed in this article are eligibility traces, reinforcement signals, reward-dependent synaptic release probability, reward-modulated STDP, and actor-critic use of TD error. The metaphor helps frame the learning question; it does not replace the variables and synaptic mechanisms used in those proposals.
Common Mistakes
Treating an eligibility trace as the reinforcement signal itself.
The trace marks an earlier synaptic event, while Miller's critic-like sensory analyzer provides reinforcement signals to neurons.
Fix:
Keep the roles distinct: eligibility identifies a synapse that can receive later credit, and the analyzer supplies the reinforcement-related signal.Assuming that a later reward automatically changes every synapse.
Miller's rule is described as using synaptically local contingent eligibility traces.
Fix:
Ask which synapse was made eligible by the earlier event before assigning credit.Confusing release probability with a single release event.
The hedonistic synapse changes its neurotransmitter release probability according to whether reward follows release or failure to release.
Fix:
Separate the event that occurred from the learned probability governing future release.Claiming that Miller's proposal and actor-critic methods are identical.
The source presents Miller's proposal as anticipating an important actor-critic feature and connects later actor-critic methods with TD error.
Fix:
Describe them as historically and conceptually connected through the problem of linking past events to later reinforcement.Treating bacterial chemotaxis as the neural learning rule.
The metaphor and the learning mechanisms are different levels of explanation.
Fix:
Use chemotaxis only to frame the consequence-guided learning problem, then return to the neural variables.
Check Your Understanding
A synapse is active during an earlier sensory event. A reinforcement signal arrives later. Explain why a contingent eligibility trace is useful, identify the role of the critic-like sensory analyzer, and state what variable changes in the hedonistic-synapse proposal.
Hints
- Start with the time gap between the earlier synaptic event and the later reinforcement signal.
- Distinguish the local eligibility trace from the analyzer's reinforcement signal.
- Name neurotransmitter release probability as the hedonistic synapse's changing variable.
What do you think happens?
Which description best separates the eligibility trace from the critic-like analyzer?
Reveal answer
Answer: The trace marks an earlier synaptic event, while the analyzer supplies a reinforcement-related signal.
Miller's proposal gives the trace and the critic-like sensory analyzer different roles: local eligibility connects an earlier event to later credit, while the analyzer provides reinforcement signals to neurons.
The Mechanism in One Pass
- An earlier sensory or synaptic event occurs.
- A synaptically local contingent eligibility trace makes the relevant synapse temporarily eligible.
- Miller's critic-like sensory analyzer supplies a reinforcement signal to neurons.
- The later reinforcement signal can affect the synapse whose earlier activity left it eligible.
- In the hedonistic-synapse proposal, reward changes neurotransmitter release probability according to whether reward follows release or failure to release.
- Reward-modulated STDP and actor-critic methods, including the use of TD error, are connected by the same broad problem of assigning later reinforcement to earlier events.
The essential idea is delayed credit assignment: eligibility preserves a local connection to the past, and reinforcement supplies information about the later consequence.
Key Takeaways
- Miller's 1981 Law-of-Effect-like rule used synaptically local contingent eligibility traces to connect earlier synaptic events with later reinforcement.
- Miller's sensory analyzer had a critic-like role because it supplied reinforcement signals to neurons.
- A hedonistic synapse changes neurotransmitter release probability according to whether reward follows release or failure to release.
- Reward-modulated STDP and actor-critic methods are related to this historical work through the problem of linking past events with later reinforcement; actor-critic methods use TD error.
- Bacterial chemotaxis is a metaphor for consequence-guided learning, not the neural learning mechanism itself.