Concepts / Actor-Critic Learning Rules

Actor-Critic Learning Rules

An actor-critic neural network contains distinct critic and actor components.

  • Programming

Two Jobs in One Learner

An actor-critic neural network divides learning into two connected jobs. The critic evaluates the agent's situation by producing state values and a temporal-difference error, or TD error. The actor uses that shared learning signal to adjust its policy, which determines the action output. The central idea is not merely that two components exist. They receive related state information and learn through one connected signal.

state informationstate informationproducesguides policy learningState featuresφ1 through φnCriticvalue and TD errorTD errorshared reinforcement signalActoraction vector
What does the critic produce, what does the actor produce, and how are their roles connected during learning?

Keep the roles separate: the critic evaluates states and produces the TD error; the actor produces the action vector and adjusts its policy using that TD error.

Routing State Features

Begin with the agent's view of its environment. This view is represented by state features written as φ1, φ2, through φn. The same collection of features is supplied to both parts of the network. One set of weighted connections leads to the critic's value unit V. Another set leads to actor units A1 through Ak. The critic weights parameterize the value function, while the actor weights parameterize the policy.

weighted connectionsmaps toweighted connectionsmaps toφ1 through φnstate featuresCritic weightsvalue parametersVstate valueActor weightspolicy parametersA1 through Akaction vector
How do the shared state features reach the critic's value output and the actor's action outputs?

Tracing One State Through Both Components

Trace state features φ1 through φn through an actor-critic network.

Start with the state: The agent's situation is represented by the collection of state features φ1 through φn.

Follow the critic path: Weighted connections carry those features to the critic's value unit V. The critic's weights determine how the state is valued.

Follow the actor path: A separate set of weighted connections carries the same features to actor units A1 through Ak. Together, these outputs form the action vector.

Separate the outputs: The value output evaluates the state, while the action vector is the actor's policy output. The two paths begin with related state information but have different jobs.

The same state features reach both components, but critic weights produce state values and actor weights produce the action vector.

Building the TD Error

The critic does more than report a value. It processes reward together with information about the current change in its estimated state values. These ingredients are combined to produce the TD error. This makes the TD error the critic's evaluation signal for learning: it reflects reward and how the critic's state-value estimates are changing.

participatesestimated valuevalue informationproducesRewardenvironment signalCriticcombines learninginformationTD errorshared reinforcement signalCurrent valuecritic estimateValue changeestimated state values
How are reward and estimated state values combined to produce the TD error?

A Qualitative TD Error Trace

Describe the critic's processing after the agent observes a state, receives a reward, and has estimated state-value information.

Observe the learning information: The critic has reward information and information about the current change in its estimated state values.

Combine the information: The critic's circuitry combines the reward signal with the value-change information.

Create the shared signal: The resulting TD error becomes the reinforcement signal used to guide learning in both components.

The critic transforms reward and estimated value-change information into one TD error that can update value and policy parameters.

One Signal, Two Updates

The TD error is shared, but its effect depends on which parameters it reaches. On the critic side, it changes the state-value parameters, so future evaluations of states can change. On the actor side, it changes the policy parameters, so the mapping from state features to the action output can change. The same signal therefore produces two different learning effects.

updateschangesupdateschangesTD errorshared signalValue parameterscritic weightsState valueshow states are valuedPolicy parametersactor weightsAction mappingstate to action output
How does one TD error signal flow to the critic's value parameters and the actor's policy parameters, and what changes in each network?
processedvalue informationcombined by criticupdatesupdatesState featuresagent situationValues and actioncritic and actor outputsTD errorcritic-generated signalCritic updatevalue parametersRewardlearning informationActor updatepolicy parameters
What happens first after state, action, and reward information are observed, and how does the learning signal reach both updates?
ComponentReceives or usesParameters changedMeaning of the update
CriticState features, reward, and estimated state-value informationState-value parametersChanges how states are valued
ActorState features and the TD error as guidancePolicy parametersChanges how state features map to the action output

Mistakes in Signal Tracing

  • Treating the actor and critic as if they had the same output.

    The critic produces state values and the TD error, whereas the actor produces the action vector.

    Fix: Name the critic as the evaluator and the actor as the policy-producing component.

  • Sending reward directly into the actor.

    The critic processes reward with estimated state-value information to produce the TD error.

    Fix: Draw reward into the critic's TD-error production, then draw the TD error to the actor.

  • Assuming that only the critic is updated.

    The TD error is the shared reinforcement signal that changes both value and policy parameters.

    Fix: Trace two update paths from the TD error: one to critic value parameters and one to actor policy parameters.

  • Forgetting that both components receive related state information.

    The same collection of state features is supplied to both parts through weighted connections.

    Fix: Route the state features separately to the critic's value unit and the actor units.

When explaining an actor-critic diagram, trace arrows in this order: state features to both components, critic processing to the TD error, and TD error to both parameter updates. This order keeps the information paths distinct.

Check the Learning Path

MEDIUM

A diagram shows state features entering a critic value unit and actor units. Reward enters the critic, which produces a TD error. The TD error then reaches two sets of weights. Identify what each set of weights represents and what changes when it is updated.

Hints
  • One set belongs to the critic and determines the value function.
  • The other set belongs to the actor and determines the policy.
  • State the different meaning of changing each set.

What do you think happens?

If the TD error is removed from the diagram after the critic produces it, which learning path is missing?

  • Only the critic's value-parameter update
  • Only the actor's policy-parameter update
  • Both the critic's and actor's updates
  • Neither update
Reveal answer

Answer: Both the critic's and actor's updates

The TD error is the shared reinforcement signal. It changes the critic's value parameters and the actor's policy parameters.

Learning Rule Summary

  1. The critic evaluates states, produces state values, and creates the TD error.
  2. The actor uses state features to produce a k-dimensional action vector through its policy parameters.
  3. The same state features are supplied to both components through separate weighted connections.
  4. The critic combines reward with information about changes in estimated state values to produce the TD error.
  5. The TD error updates critic value parameters and actor policy parameters, giving the two updates different meanings.

Key Takeaways

  • The critic evaluates states and produces the TD error.
  • The actor produces the action vector and adjusts its policy.
  • State features reach both components through weighted connections.
  • Reward contributes to the critic's TD-error calculation rather than reaching the actor through a direct path.
  • The shared TD error updates value parameters in the critic and policy parameters in the actor.