Concepts / Reward-Modulated Synaptic Plasticity

Reward-Modulated Synaptic Plasticity

Eligibility traces connect earlier neural activity with a later learning consequence.

  • Programming

The Delayed-Credit Problem

A learning signal often arrives after the neural activity that should receive credit. A synapse may therefore need to retain information about an earlier activity event until a later reward or other learning consequence appears. An eligibility trace provides this bridge. It marks activity as relevant for a possible later change; it does not mean that the synapse must change immediately.

leaves informationremains relevant untilcan help determineNeural activityearlier eventEligibility traceactivity retainedLearning consequencelater signalSynaptic changepossible update
How does an earlier synaptic activity event remain available until a later reward or consequence arrives?

The central timing idea is simple: activity happens first, eligibility preserves its relevance, and a later learning signal can then use that retained information.

Two Actor Traces

The history of actor-critic learning includes two forms of an actor-unit trace. An early form was written as A × φ(s). A later form was written as (A - π(A|S,θ)) × φ(s). These expressions show a historical change in what the actor trace records: the later form includes a policy-related term rather than only the product of the actor quantity and the state representation.

Trace formWhen it appears in the historical accountWhat it recordsRole in the learning discussion
A × φ(s)Early actor-unit traceThe product of the actor quantity A and the state representation φ(s)An earlier way of representing activity relevant to the actor
(A - π(A|S,θ)) × φ(s)Later policy-related traceThe difference between A and the policy term π(A|S,θ), combined with φ(s)A later form that relates the trace to the policy

The two expressions are historically related actor-unit traces, but they are not the same expression.

representsincludesA × φ(s)early actor-unit trace(A - π(A|S,θ)) ×φ(s)later policy-related traceActor activityactivity representationPolicy relationactivity relative to π
What does each trace record, when does it appear, and how do its timing and role differ?

Activity Across Levels

Eligibility can be described at several levels. Synaptically local eligibility is associated with an individual synapse. Neuronal eligibility is associated with activity at the level of an entire neuron. Computational actor-unit traces describe a related idea in an actor-critic learning rule. These are not merely different names for exactly the same object: they identify different levels at which information about activity and possible consequences is retained.

This distinction also appears in the historical development of the ideas. Klopf's ideas about local synaptic eligibility inspired the actor-critic algorithm implemented by Barto, Sutton, and Anderson as an artificial neural network with a single neuron-like actor unit. Crow proposed a related but different form in 1968, in which contingent eligibility was associated with entire neurons rather than individual synapses.

can be discussed ascan be discussed ascan be represented asEligibilityrelated ideaSynaptic eligibilityindividual synapseNeuronal eligibilityentire neuronActor-unit tracecomputational learning rule
At what level does the system retain information about activity and its possible consequences?

One Event, Three Descriptions

An earlier neural activity event is followed later by a learning consequence. How can the same general eligibility idea be described at different levels?

Synaptic level: Describe the retained information as eligibility associated with an individual synapse.

Neuronal level: Describe the retained information as contingent eligibility associated with an entire neuron.

Computational level: Describe the retained information as an actor-unit trace used in an actor-critic learning rule.

The descriptions are related because each connects earlier activity with a later consequence, but they differ in the level at which the retained information is represented.

Actor-Critic Information Flow

In an actor-critic learning rule, neural activity contributes to an actor-unit trace. The critic supplies a prediction-related part of the learning process, while a reward or other later learning signal supplies the consequence that makes earlier activity relevant. The trace allows the earlier activity to remain available for the learning update instead of requiring the consequence to arrive at the same instant.

contributes tocarries earlier activitycontributes predictionprovides consequenceNeural activityearlier eventActor-unit traceretained activityCritic predictionpredictionRewardlater signalLearning updatepossible change
How do neural activity, the actor's trace, the critic's prediction, and the reward signal flow through the learning update?

What do you think happens?

An actor-related neural event occurs first, and a reward arrives later. What is the eligibility trace doing between those two events?

  • It identifies earlier activity that may be relevant to the later learning signal
  • It guarantees that the synapse changes immediately
  • It replaces the critic's prediction
  • It prevents the later reward from affecting learning
Reveal answer

Answer: It identifies earlier activity that may be relevant to the later learning signal.

Eligibility retains information about activity so that a later consequence can help determine whether that activity should contribute to a change. It does not itself guarantee an immediate synaptic change.

Plasticity and STDP

The history of eligibility traces overlaps with research on reward-modulated synaptic plasticity. Reynolds and Wickens proposed a three-factor rule for plasticity in the corticostriatal pathway: dopamine modulates changes in corticostriatal synaptic efficacy. This comparison adds a modulatory factor to the activity-related factors involved in synaptic change.

STDP, or spike-timing-dependent plasticity, concerns synaptic plasticity that depends on the relative timing of pre- and postsynaptic spikes. Earlier evidence showed that the timing relationship between spikes matters for changing synaptic efficacy, and the definitive demonstration is attributed to Markram, Lübke, Frotscher, and Sakmann in 1997.

The relationship between STDP, eligibility, and temporal-difference ideas has not been treated as one settled equivalence. Rao and Sejnowski suggested that STDP could arise from a TD-like mechanism at synapses with non-contingent eligibility traces lasting about 10 milliseconds. Dayan later commented that this proposal would require an error like the one in Sutton and Barto's early model of classical conditioning rather than a true TD error. These remarks show that the connections have been actively debated.

Dopamine also appears directly in later work connecting reward modulation and spike timing. Pawlak and Kerr reported that dopamine is necessary to induce STDP at corticostriatal synapses of medium spiny neurons. Together with work on reward-modulated STDP, this places dopamine and spike timing within the broader discussion of how synaptic changes may depend on both activity and later signals.

historical developmentrelated discussionbroadens connectionCrow1968 neuronal eligibilitySTDP1997 definitivedemonstrationThree-factorplasticitydopamine modulationDopamine and STDPcorticostriatal synapses
How did ideas about eligibility, STDP, and reward modulation develop and connect over time?

Common Interpretation Errors

  • Treating an eligibility trace as an immediate synaptic update

    Eligibility marks activity as relevant if a later learning signal arrives; it does not mean that a change happens immediately.

    Fix: Separate the earlier activity event, the retained eligibility information, and the later learning consequence.

  • Treating the early and later actor traces as identical

    The source identifies them as historically different forms, and the later expression includes a policy-related term.

    Fix: Name which form is being discussed and note whether the policy term appears.

  • Assuming synaptic and neuronal eligibility refer to the same level

    The distinction concerns where the system retains information about activity and possible consequences.

    Fix: Use synaptic eligibility for the individual-synapse level and neuronal eligibility for the entire-neuron level.

  • Claiming that STDP and a TD error are automatically equivalent

    The source describes proposals connecting STDP with TD-like mechanisms and also records a later criticism of that connection.

    Fix: Describe the relationship as a historical proposal and debate, not as a settled equivalence.

Practice the Distinction

MEDIUM

A learning system records an earlier actor-related event and receives a reward later. Explain, in three parts, what the eligibility trace contributes, which historical actor-unit trace form you are using, and whether the description is synaptic, neuronal, or computational.

Hints
  • Begin by explaining why the earlier activity must remain available.
  • Compare A × φ(s) with (A - π(A|S,θ)) × φ(s).
  • State the level at which the retained information is being described.

Classifying a Delayed-Reward Scenario

An artificial actor unit is active before a later reward. The explanation uses (A - π(A|S,θ)) × φ(s). Classify the trace and explain its timing role.

Identify the level: Because the expression is an actor-unit expression in an actor-critic learning rule, this is a computational actor-unit trace.

Identify the historical form: The expression is the later policy-related form because it contains the policy term π(A|S,θ).

Identify the timing role: The trace retains information about the earlier actor-related activity so that the later reward can make that activity relevant to a learning update.

Avoid overclaiming: The trace does not by itself prove that a synapse changes immediately, and the example does not establish that this computational trace is identical to a biological synaptic trace.

This is a later policy-related computational actor-unit trace that bridges earlier activity and a later reward without equating computational, neuronal, and synaptic descriptions.

Key Takeaways

  1. Eligibility traces retain information about earlier neural activity so a later learning consequence can use it.
  2. The early actor-unit trace was written as A × φ(s), while the later policy-related trace was written as (A - π(A|S,θ)) × φ(s).
  3. Synaptic, neuronal, and computational actor-unit eligibility describe related ideas at different levels.
  4. Reward-modulated plasticity connects activity-related factors with dopamine or another later modulatory signal.
  5. STDP emphasizes relative pre- and postsynaptic spike timing, and its relationship with eligibility and temporal-difference learning has been historically debated.

Key Takeaways

  • Eligibility traces solve a timing problem by retaining the relevance of earlier activity until a later consequence arrives.
  • Actor-unit traces changed historically from A × φ(s) to (A - π(A|S,θ)) × φ(s), adding a policy-related term.
  • Eligibility can be described locally at synapses, across neurons, or computationally in an actor-critic rule.
  • Reward modulation, dopamine, and STDP broaden the discussion of how activity and later signals can influence synaptic efficacy.
  • The historical relationship among eligibility, STDP, and temporal-difference learning includes proposals and debate rather than one settled equivalence.