Neuromodulatory Systems
The hedonistic neuron hypothesis applies reinforcement learning at the level of an individual neuron.
A Trainable Single Neuron
Many descriptions of reinforcement learning begin with an animal, an agent, or a larger neural network. Klopf's hedonistic neuron hypothesis starts at a smaller scale: one neuron. The proposal is that an individual neuron can change the efficacies of its synapses according to what follows its own action potentials. In this view, a neuron is not merely receiving fixed inputs. It can be trained through response-contingent reinforcement, in a way related to instrumental conditioning.
The Eligibility Trace
A central problem is credit assignment within the neuron: if several synapses are present, which ones should be changed when a consequence follows the neuron's action potential? The hypothesis points to local molecular traces at synapses. A trace associated with a synapse can help determine whether that synapse is eligible for a later modification. The important idea is locality: the information relevant to eligibility is held at or near the synapse rather than being treated as an undifferentiated change to every synapse.
Reward and Punishment
Klopf's proposal is response-contingent: the consequence follows the neuron's own action potentials. Synaptic efficacies are proposed to change according to whether that consequence is rewarding or punishing. Inputs treated as rewarding should become more influential relative to inputs treated as punishing. The hypothesis therefore connects three events: the neuron's action, what follows that action, and a later change in the influence of synapses.
A hypothetical consequence-dependent update
Suppose a neuron receives two inputs. One input is treated as rewarding when it contributes around an action potential, and another is treated as punishing. What directional change does the hypothesis predict?
Action: The neuron produces an action potential after receiving its inputs.
Consequence: The action is followed by a rewarding or punishing consequence.
Eligibility: Local molecular traces help determine which synapses are eligible for a later modification.
Relative influence: The input treated as rewarding should become more influential relative to the input treated as punishing.
The proposed learning rule changes synaptic efficacies in relation to the consequence of the neuron's own action potential; it does not simply change all synapses in the same direction.
The Reinforcement Pathway
In the original formulation described here, reinforcing information reaches the neuron through synaptic input. That input can be part of the same general synaptic influence that excites or inhibits the neuron's spike-generating activity. Klopf also wanted to avoid depending on a centralized source of training information.
The source also describes a possible historical update. If Klopf had known what is now known about neuromodulatory systems, he might have assigned the reinforcing role to neuromodulatory input. Neuromodulatory input is therefore a possible way to think about how reinforcing signals could reach neurons, while the original hypothesis used synaptic input for that role.
Common Misreadings
Treating the hypothesis as a rule for an entire animal or centralized controller.
Klopf's hypothesis starts with a single neuron and aims to avoid dependence on a centralized source of training information.
Fix:
Describe the individual neuron as changing its synaptic efficacies according to what follows its own action potentials.Assuming that reinforcement changes every synapse identically.
Local molecular traces help determine which synapses are eligible for later modification.
Fix:
Keep the local eligibility step in the explanation: the consequence affects synapses selected as eligible by their local traces.Presenting neuromodulatory input as the original formulation.
The original formulation described reinforcing information as arriving through synaptic input.
Fix:
State that neuromodulatory input is a possible alternative role discussed in light of knowledge about neuromodulatory systems.Confusing an eligibility trace with the consequence itself.
The source assigns the trace a role in determining eligibility, while reinforcing information is a separate part of the proposed process.
Fix:
Explain the trace as local information that helps identify a synapse for later modification.
Check Your Understanding
Explain the proposed sequence in your own words: a synaptic input contributes to a neuron's action potential, a consequence follows, a local molecular trace helps identify an eligible synapse, and the consequence is associated with a later change in synaptic efficacy. Then state how the explanation differs between the original synaptic-input formulation and the possible neuromodulatory alternative.
Hints
- Begin with the individual neuron rather than an entire animal or centralized instructor.
- Distinguish local eligibility from the reinforcing information that arrives later.
- Use possible or proposed when describing the neuromodulatory role.
Key Takeaways
- Klopf's hedonistic neuron hypothesis applies reinforcement learning at the level of an individual neuron.
- A neuron's synaptic efficacies are proposed to change according to rewarding or punishing consequences of its action potentials.
- Local molecular traces help determine which synapses are eligible for a later modification.
- The original formulation used synaptic input to convey reinforcing information, while neuromodulatory input is discussed as a possible alternative role.
- Inputs treated as rewarding should become more influential relative to inputs treated as punishing.
Key Takeaways
- The hedonistic neuron hypothesis places reinforcement learning within a single neuron.
- Rewarding and punishing consequences are proposed to alter the relative influence of synaptic inputs.
- Synaptically local molecular traces help identify which synapses can be modified later.
- Reinforcement was originally described as synaptic input, with neuromodulatory input offered as a possible alternative route.