Synaptic Plasticity and Reinforcement Signals
Miller's 1981 Law-of-Effect-like rule included synaptically local contingent eligibility traces.
The Delayed-Credit Problem
A central problem in reinforcement learning is temporal credit assignment: how can a system connect a past synaptic event with a reinforcement signal that arrives later? Miller's 1981 proposal addressed this problem with a Law-of-Effect-like learning rule containing synaptically local contingent eligibility traces. The trace preserves the relevance of a recent synaptic event until reinforcement can determine whether learning should occur.
The important idea is not that reinforcement must arrive at the same moment as synaptic activity. Instead, a local eligibility trace keeps a recent event available for later evaluation.
Miller's Learning Rule
Miller's rule was Law-of-Effect-like: synaptic learning was related to a later reinforcement signal, rather than being determined only by the local synaptic event itself. Its distinctive mechanism was the contingent eligibility trace. The trace was synaptically local, meaning that the relevant record of recent activity was associated with the synapse where the event occurred. When reinforcement arrived, that stored eligibility could participate in determining whether the connection changed.
A Delayed Reinforcement Episode
A synaptic event occurs first, and a reinforcement signal arrives later. How does Miller's proposal connect them?
Event: The synapse undergoes a relevant recent activity event.
Local trace: The synapse stores a contingent eligibility trace associated with that recent event.
Later reinforcement: A reinforcement signal arrives after the synaptic event rather than simultaneously with it.
Learning decision: The later signal can use the stored eligibility to determine whether the earlier synaptic event should contribute to synaptic change.
The eligibility trace supplies a bridge between past local synaptic activity and later reinforcement.
The Critic-Like Sensory Analyzer
Miller's proposal also included a sensory analyzer unit with a critic-like role. The source characterizes this unit as a mechanism that provided reinforcement signals to neurons. This adds an evaluating component to the learning system: synaptic activity is not the only relevant process, because another mechanism supplies information that can reinforce or evaluate what happened.
Critic-like mechanism: a mechanism that evaluates or supplies reinforcement information to neurons, rather than being described merely as the synapse that undergoes change.
The Hedonistic Synapse
Seung's hedonistic synapse focuses on an individual synapse and its neurotransmitter release probability. The synapse changes that probability according to whether reward follows release or follows failure to release. In this view, the synapse is sensitive not only to its own release event but also to the reward relationship that follows it.
Connections to Later Frameworks
The historical proposals can be understood as addressing the same broad coordination problem later seen in reward-modulated STDP and actor-critic methods: how should a system connect a past synaptic event with a later reinforcement signal? The source connects Miller's proposal with general features of reward-modulated STDP and connects the later actor-critic architecture with TD error. These are historical and conceptual relationships, not a claim that all of the mechanisms are identical.
| Idea | Role in the source pack |
|---|---|
| Miller's proposal | Uses a Law-of-Effect-like rule with synaptically local contingent eligibility traces. |
| Reward-modulated STDP | Is connected by the source to general features of Miller's proposal. |
| Actor-critic methods | Later use a critic and TD error as part of the historical refinement. |
| Central question | How can a past synaptic event be connected with a later reinforcement signal? |
Metaphor and Mechanism
Bacterial chemotaxis belongs in this topic as a metaphor or source of inspiration, not as the synaptic learning mechanism itself. The learning mechanisms discussed here are the eligibility trace, the critic-like reinforcement signal, and the reward-dependent change in neurotransmitter release probability. Keeping these levels separate prevents a biological analogy from being mistaken for a literal description of Miller's rule, the hedonistic synapse, reward-modulated STDP, or actor-critic learning.
When explaining the analogy, name the abstract learning operation explicitly. Say that chemotaxis is a metaphor or inspiration, then separately identify the trace, reinforcement signal, or release-probability update being discussed.
Mistakes in Reasoning
Treating the eligibility trace as the reinforcement signal.
Miller's proposal separates the local record of recent synaptic activity from the later reinforcement signal.
Fix:
Describe the trace as preserved eligibility and the critic-like mechanism as a source of reinforcement information.Assuming that all three frameworks are identical.
The source connects them historically and conceptually, but describes them as related ideas and later refinements.
Fix:
Use the shared credit-assignment question to relate them while preserving their distinct names and roles.Ignoring whether reward followed release or failure to release.
The source emphasizes the relationship between reward and the preceding release outcome.
Fix:
State whether the relevant preceding event was neurotransmitter release or failure to release.Turning the chemotaxis metaphor into a literal synaptic mechanism.
The metaphor and the inspired learning mechanism are different explanatory levels.
Fix:
Separate the metaphor from the mechanisms explicitly named in the source.
Practice the Credit Assignment
A synaptic event occurs, a local eligibility trace is retained, and reinforcement arrives later. Identify which part records the past event, which part supplies later evaluation, and which part is ultimately influenced by the relationship between the two.
Hints
- Look for the mechanism described as synaptically local.
- Look for the critic-like mechanism that provides reinforcement signals.
- The final learning target is the connection or synapse.
Compare these two descriptions: one says that reward follows neurotransmitter release; the other says that reward follows failure to release. Explain why the hedonistic synapse treats them as different cases, and identify the synaptic quantity that changes.
Hints
- The source distinguishes release from failure to release.
- The changed quantity concerns neurotransmitter release.
Summary
- Miller's 1981 Law-of-Effect-like rule used a synaptically local contingent eligibility trace to connect past synaptic activity with later reinforcement.
- Miller's sensory analyzer was critic-like because it supplied reinforcement signals to neurons.
- Seung's hedonistic synapse changed an individual synapse's neurotransmitter release probability according to whether reward followed release or failure to release.
- The source relates these historical ideas to reward-modulated STDP and to the later actor-critic use of TD error.
- Bacterial chemotaxis should be treated as a metaphor or inspiration, not confused with the synaptic learning mechanisms themselves.
Key Takeaways
- Eligibility traces solve a delayed-credit problem by preserving recent local synaptic activity until reinforcement arrives.
- Miller's critic-like sensory analyzer supplies reinforcement information that can influence neuronal learning.
- The hedonistic synapse links reward to changes in neurotransmitter release probability, distinguishing release from failure to release.
- Reward-modulated STDP and actor-critic methods are later frameworks connected to the same broad question of relating past events to later reinforcement.
- A chemotaxis metaphor can motivate an idea without being identical to the learning mechanism it inspired.