Policy-Gradient Actor-Critic Method
The critic estimates state value by summing weighted state features.
Why Timing Matters
A critic estimates how valuable the current state is, but the reinforcement that should change the critic may arrive later. The state that helped produce that reinforcement may no longer be the current state when the signal appears. The critic addresses this temporal credit-assignment problem by giving each critic synapse a temporary memory of recent presynaptic activity. This memory is called an eligibility trace.
The trace does not itself provide the reinforcement signal. It preserves information about which synapses recently received feature activity, so a later signal can modify the appropriate connections.
Building a State-Value Estimate
The critic starts with a feature-vector representation of the current state. Each state feature supplies an input to a simple linear, neuron-like unit. The unit multiplies each feature by its corresponding critic weight and adds the weighted inputs together. The resulting output is the critic's estimate of the value of the current state.
Combining Three State Features
A current state is represented by three features. The critic has one corresponding weight for each feature.
Represent the state: The critic receives the feature-vector representation of the current state. Each feature is an input to the linear critic unit.
Apply the synaptic weights: Each feature is combined with its corresponding weight. The weight determines how strongly that feature contributes to the critic's output.
Add the contributions: The critic adds the weighted feature contributions together. This total is the estimate of the value of the current state.
The critic represents state value as a weighted sum of state features.
Reading a Synapse Trace
Each critic synapse keeps its own fading record. When a feature from the current state reaches that synapse, the activity adds to the synapse's eligibility trace. Afterward, the trace continually decays. If a later state produces activity at the same synapse, that activity contributes to the existing trace as well. While the trace remains nonzero, the synapse is eligible for modification.
A Trace That Fades Across States
A feature is active in one state, and reinforcement arrives after the system has moved through another state.
Initial activity: The feature signal reaches its critic synapse. That activity creates a nonzero eligibility trace at the synapse.
Decay: As time passes, the trace fades. It remains a record of the earlier feature activity while it is nonzero.
Additional activity: If another state activates the same feature, that activity contributes to the trace at the same synapse.
Reinforcement: When the reinforcement signal arrives, the synapse's remaining trace indicates how strongly it is positioned to be modified.
The trace links earlier presynaptic activity with a later reinforcement signal.
Non-Contingent Trace Formation
The critic's traces are called non-contingent because they depend on presynaptic activity only. Presynaptic activity is the feature signal arriving at the synapse. Activity produced by the critic unit on the postsynaptic side is not required to create the trace. This is the key distinction from a trace whose formation depends on both presynaptic and postsynaptic activity.
Critic and Reinforcement
The critic's weight update combines three ingredients: a learning-rate parameter, the reinforcement signal, and the eligibility-trace vector. In the source rule, the weight vector changes by adding the product of the learning-rate parameter β, the reinforcement signal δ, and the eligibility vector ew. Conceptually, the reinforcement signal is broadcast to the critic's synapses, while each synapse's trace determines how strongly that synapse is positioned to be modified.
When tracing the learning rule, ask two separate questions: which synapses have a nonzero eligibility trace, and what reinforcement signal arrives while those traces remain? The first question identifies where an update can be applied; the second determines the signal driving the update.
TD Conditioning Connection
The critic learning rule connects directly with the temporal-difference model of classical conditioning. In this connection, the critic's temporal-difference error corresponds to the prediction-error signal used to explain learning in classical conditioning. The source describes the critic's learning rule as essentially that model, and notes that its TD errors parallel dopamine neuron activity.
| Critic perspective | TD conditioning perspective |
|---|---|
| The critic estimates the value of a state. | Learning is explained through a prediction-error signal. |
| The critic uses an eligibility-trace vector with the reinforcement signal. | The temporal-difference error parallels the signal used in the conditioning model. |
| The update changes critic weights associated with eligible synapses. | The learning-rule connection parallels dopamine neuron activity. |
Common Misreadings
Treating the critic output as a direct record of reinforcement.
The critic output is a state-value estimate produced by summing weighted state features.
Fix:
Separate the value estimate from the later reinforcement signal used in the weight update.Assuming that only the currently active state can be credited.
Eligibility traces preserve a decaying record of past presynaptic activity.
Fix:
Check whether the feature's synapse still has a nonzero trace when reinforcement arrives.Assuming that reinforcement creates the eligibility trace.
The critic's non-contingent trace is built from presynaptic feature activity and then decays.
Fix:
Describe reinforcement as part of the later update, not as the source of the trace.Calling a trace non-contingent because it ignores reinforcement entirely.
Non-contingent describes how the trace is formed, not whether the later update uses reinforcement.
Fix:
State that trace formation depends on presynaptic activity only, while the later update combines the trace with reinforcement.Equating the critic with every biological detail of a dopamine system.
The source describes a relationship between the learning rule, TD errors, and dopamine neuron activity, not complete biological identity.
Fix:
Use the connection to explain the prediction-error relationship and keep the biological claim limited.
Check Your Understanding
A feature is active in an earlier state. Its critic synapse then has a fading, nonzero eligibility trace when a reinforcement signal arrives. Explain what each of the following contributes: the state feature, the eligibility trace, and the reinforcement signal. Then explain why the trace is called non-contingent.
Hints
- Start with the feature's role as a presynaptic input to the critic synapse.
- Describe the trace as a temporary record that decays over time.
- Describe reinforcement as part of the later weight update.
- For non-contingent, identify which side of the synapse is required to form the trace.
What do you think happens?
A feature was active earlier, but its eligibility trace has completely decayed before reinforcement arrives. Is that synapse still identified by the trace as eligible for modification?
Reveal answer
Answer: No, because the trace is no longer nonzero.
The trace is a decaying record of past presynaptic activity. A synapse is described as eligible for modification while its trace remains nonzero.
Key Takeaways
- The critic estimates state value by multiplying state features by corresponding weights and adding the weighted inputs.
- Each critic synapse keeps a decaying eligibility trace so earlier feature activity can remain relevant when later reinforcement arrives.
- The critic update combines the reinforcement signal with the eligibility-trace vector, allowing the trace to identify how strongly each synapse is positioned for modification.
- Non-contingent traces depend on presynaptic feature activity and do not require postsynaptic critic-unit activity to form.
- The critic learning rule connects with the TD model of classical conditioning through its prediction-error signal and its parallel with dopamine neuron activity.
Key Takeaways
- A critic represents state value as a weighted sum of state features.
- Eligibility traces preserve a fading record of presynaptic activity at individual critic synapses.
- Later reinforcement combines with those traces to solve a timing problem in critic learning.
- Non-contingent traces are formed from presynaptic activity without requiring postsynaptic activity.
- This critic learning rule is connected to the TD model of classical conditioning and its prediction-error interpretation.