Concepts / Policy-Gradient Actor-Critic Method

Policy-Gradient Actor-Critic Method

The critic estimates state value by summing weighted state features.

  • Programming

Why Timing Matters

A critic estimates how valuable the current state is, but the reinforcement that should change the critic may arrive later. The state that helped produce that reinforcement may no longer be the current state when the signal appears. The critic addresses this temporal credit-assignment problem by giving each critic synapse a temporary memory of recent presynaptic activity. This memory is called an eligibility trace.

adds activitytime passesidentifies eligibilitydrives updateState featurepresynaptic activityEligibility tracenonzeroEligibility tracedecayingReinforcement signallater signalCritic synapseeligible for modification
How does earlier feature activity remain available when a later reinforcement signal arrives?

The trace does not itself provide the reinforcement signal. It preserves information about which synapses recently received feature activity, so a later signal can modify the appropriate connections.

Building a State-Value Estimate

The critic starts with a feature-vector representation of the current state. Each state feature supplies an input to a simple linear, neuron-like unit. The unit multiplies each feature by its corresponding critic weight and adds the weighted inputs together. The resulting output is the critic's estimate of the value of the current state.

through w1through w2through wjmultiply and addproducesFeature φ1(s)inputCritic weightsw1, w2, …, wjWeighted sumvalue estimateState valuecritic outputFeature φ2(s)inputFeature φj(s)input
How are several state features combined through critic synapses to estimate the value of the current state?

Combining Three State Features

A current state is represented by three features. The critic has one corresponding weight for each feature.

Represent the state: The critic receives the feature-vector representation of the current state. Each feature is an input to the linear critic unit.

Apply the synaptic weights: Each feature is combined with its corresponding weight. The weight determines how strongly that feature contributes to the critic's output.

Add the contributions: The critic adds the weighted feature contributions together. This total is the estimate of the value of the current state.

The critic represents state value as a weighted sum of state features.

Reading a Synapse Trace

Each critic synapse keeps its own fading record. When a feature from the current state reaches that synapse, the activity adds to the synapse's eligibility trace. Afterward, the trace continually decays. If a later state produces activity at the same synapse, that activity contributes to the existing trace as well. While the trace remains nonzero, the synapse is eligible for modification.

retains eligibilityretains eligibilityretains eligibilitycombines with tracesSynapse 1feature input + fadingtraceLater reinforcementbroadcast signalWeight updatetrace selects strengthSynapse 2feature input + fadingtraceSynapse jfeature input + fadingtrace
What information is retained at each critic synapse between feature activity and a later update?

A Trace That Fades Across States

A feature is active in one state, and reinforcement arrives after the system has moved through another state.

Initial activity: The feature signal reaches its critic synapse. That activity creates a nonzero eligibility trace at the synapse.

Decay: As time passes, the trace fades. It remains a record of the earlier feature activity while it is nonzero.

Additional activity: If another state activates the same feature, that activity contributes to the trace at the same synapse.

Reinforcement: When the reinforcement signal arrives, the synapse's remaining trace indicates how strongly it is positioned to be modified.

The trace links earlier presynaptic activity with a later reinforcement signal.

Non-Contingent Trace Formation

alone is sufficientone required inputsecond required inputPresynapticactivityfeature signalEligibility tracetrace formsPresynapticactivityfeature signalEligibility tracerequires both activitiesPostsynapticactivitycritic-unit activity
What is the difference between a trace triggered by presynaptic activity alone and one that also requires postsynaptic activity?

The critic's traces are called non-contingent because they depend on presynaptic activity only. Presynaptic activity is the feature signal arriving at the synapse. Activity produced by the critic unit on the postsynaptic side is not required to create the trace. This is the key distinction from a trace whose formation depends on both presynaptic and postsynaptic activity.

Critic and Reinforcement

The critic's weight update combines three ingredients: a learning-rate parameter, the reinforcement signal, and the eligibility-trace vector. In the source rule, the weight vector changes by adding the product of the learning-rate parameter β, the reinforcement signal δ, and the eligibility vector ew. Conceptually, the reinforcement signal is broadcast to the critic's synapses, while each synapse's trace determines how strongly that synapse is positioned to be modified.

provides featuresestimates valuecreates tracesparticipates in signalcombines with traceselects synapse strengthguides learningCurrent statefeature vectorCriticweighted state featuresValue estimatecritic outputCritic weightsupdatedEligibility vectorsynapse tracesReinforcement signalδActorlearning component
How do state features enter the critic, how does the critic produce a value-related signal, and how does reinforcement participate in critic learning?

When tracing the learning rule, ask two separate questions: which synapses have a nonzero eligibility trace, and what reinforcement signal arrives while those traces remain? The first question identifies where an update can be applied; the second determines the signal driving the update.

TD Conditioning Connection

The critic learning rule connects directly with the temporal-difference model of classical conditioning. In this connection, the critic's temporal-difference error corresponds to the prediction-error signal used to explain learning in classical conditioning. The source describes the critic's learning rule as essentially that model, and notes that its TD errors parallel dopamine neuron activity.

Critic perspectiveTD conditioning perspective
The critic estimates the value of a state.Learning is explained through a prediction-error signal.
The critic uses an eligibility-trace vector with the reinforcement signal.The temporal-difference error parallels the signal used in the conditioning model.
The update changes critic weights associated with eligible synapses.The learning-rule connection parallels dopamine neuron activity.

Common Misreadings

  • Treating the critic output as a direct record of reinforcement.

    The critic output is a state-value estimate produced by summing weighted state features.

    Fix: Separate the value estimate from the later reinforcement signal used in the weight update.

  • Assuming that only the currently active state can be credited.

    Eligibility traces preserve a decaying record of past presynaptic activity.

    Fix: Check whether the feature's synapse still has a nonzero trace when reinforcement arrives.

  • Assuming that reinforcement creates the eligibility trace.

    The critic's non-contingent trace is built from presynaptic feature activity and then decays.

    Fix: Describe reinforcement as part of the later update, not as the source of the trace.

  • Calling a trace non-contingent because it ignores reinforcement entirely.

    Non-contingent describes how the trace is formed, not whether the later update uses reinforcement.

    Fix: State that trace formation depends on presynaptic activity only, while the later update combines the trace with reinforcement.

  • Equating the critic with every biological detail of a dopamine system.

    The source describes a relationship between the learning rule, TD errors, and dopamine neuron activity, not complete biological identity.

    Fix: Use the connection to explain the prediction-error relationship and keep the biological claim limited.

Check Your Understanding

MEDIUM

A feature is active in an earlier state. Its critic synapse then has a fading, nonzero eligibility trace when a reinforcement signal arrives. Explain what each of the following contributes: the state feature, the eligibility trace, and the reinforcement signal. Then explain why the trace is called non-contingent.

Hints
  • Start with the feature's role as a presynaptic input to the critic synapse.
  • Describe the trace as a temporary record that decays over time.
  • Describe reinforcement as part of the later weight update.
  • For non-contingent, identify which side of the synapse is required to form the trace.

What do you think happens?

A feature was active earlier, but its eligibility trace has completely decayed before reinforcement arrives. Is that synapse still identified by the trace as eligible for modification?

  • Yes, because any past activity remains permanently available
  • No, because the trace is no longer nonzero
  • Yes, but only if the critic output was active
  • No, because feature activity never creates a trace
Reveal answer

Answer: No, because the trace is no longer nonzero.

The trace is a decaying record of past presynaptic activity. A synapse is described as eligible for modification while its trace remains nonzero.

Key Takeaways

  1. The critic estimates state value by multiplying state features by corresponding weights and adding the weighted inputs.
  2. Each critic synapse keeps a decaying eligibility trace so earlier feature activity can remain relevant when later reinforcement arrives.
  3. The critic update combines the reinforcement signal with the eligibility-trace vector, allowing the trace to identify how strongly each synapse is positioned for modification.
  4. Non-contingent traces depend on presynaptic feature activity and do not require postsynaptic critic-unit activity to form.
  5. The critic learning rule connects with the TD model of classical conditioning through its prediction-error signal and its parallel with dopamine neuron activity.

Key Takeaways

  • A critic represents state value as a weighted sum of state features.
  • Eligibility traces preserve a fading record of presynaptic activity at individual critic synapses.
  • Later reinforcement combines with those traces to solve a timing problem in critic learning.
  • Non-contingent traces are formed from presynaptic activity without requiring postsynaptic activity.
  • This critic learning rule is connected to the TD model of classical conditioning and its prediction-error interpretation.