Actor-Critic Learning Rules
An actor-critic neural network contains distinct critic and actor components.
Two Jobs in One Learner
An actor-critic neural network divides learning into two connected jobs. The critic evaluates the agent's situation by producing state values and a temporal-difference error, or TD error. The actor uses that shared learning signal to adjust its policy, which determines the action output. The central idea is not merely that two components exist. They receive related state information and learn through one connected signal.
Keep the roles separate: the critic evaluates states and produces the TD error; the actor produces the action vector and adjusts its policy using that TD error.
Routing State Features
Begin with the agent's view of its environment. This view is represented by state features written as φ1, φ2, through φn. The same collection of features is supplied to both parts of the network. One set of weighted connections leads to the critic's value unit V. Another set leads to actor units A1 through Ak. The critic weights parameterize the value function, while the actor weights parameterize the policy.
Tracing One State Through Both Components
Trace state features φ1 through φn through an actor-critic network.
Start with the state: The agent's situation is represented by the collection of state features φ1 through φn.
Follow the critic path: Weighted connections carry those features to the critic's value unit V. The critic's weights determine how the state is valued.
Follow the actor path: A separate set of weighted connections carries the same features to actor units A1 through Ak. Together, these outputs form the action vector.
Separate the outputs: The value output evaluates the state, while the action vector is the actor's policy output. The two paths begin with related state information but have different jobs.
The same state features reach both components, but critic weights produce state values and actor weights produce the action vector.
Building the TD Error
The critic does more than report a value. It processes reward together with information about the current change in its estimated state values. These ingredients are combined to produce the TD error. This makes the TD error the critic's evaluation signal for learning: it reflects reward and how the critic's state-value estimates are changing.
A Qualitative TD Error Trace
Describe the critic's processing after the agent observes a state, receives a reward, and has estimated state-value information.
Observe the learning information: The critic has reward information and information about the current change in its estimated state values.
Combine the information: The critic's circuitry combines the reward signal with the value-change information.
Create the shared signal: The resulting TD error becomes the reinforcement signal used to guide learning in both components.
The critic transforms reward and estimated value-change information into one TD error that can update value and policy parameters.
One Signal, Two Updates
The TD error is shared, but its effect depends on which parameters it reaches. On the critic side, it changes the state-value parameters, so future evaluations of states can change. On the actor side, it changes the policy parameters, so the mapping from state features to the action output can change. The same signal therefore produces two different learning effects.
| Component | Receives or uses | Parameters changed | Meaning of the update |
|---|---|---|---|
| Critic | State features, reward, and estimated state-value information | State-value parameters | Changes how states are valued |
| Actor | State features and the TD error as guidance | Policy parameters | Changes how state features map to the action output |
Mistakes in Signal Tracing
Treating the actor and critic as if they had the same output.
The critic produces state values and the TD error, whereas the actor produces the action vector.
Fix:
Name the critic as the evaluator and the actor as the policy-producing component.Sending reward directly into the actor.
The critic processes reward with estimated state-value information to produce the TD error.
Fix:
Draw reward into the critic's TD-error production, then draw the TD error to the actor.Assuming that only the critic is updated.
The TD error is the shared reinforcement signal that changes both value and policy parameters.
Fix:
Trace two update paths from the TD error: one to critic value parameters and one to actor policy parameters.Forgetting that both components receive related state information.
The same collection of state features is supplied to both parts through weighted connections.
Fix:
Route the state features separately to the critic's value unit and the actor units.
When explaining an actor-critic diagram, trace arrows in this order: state features to both components, critic processing to the TD error, and TD error to both parameter updates. This order keeps the information paths distinct.
Check the Learning Path
A diagram shows state features entering a critic value unit and actor units. Reward enters the critic, which produces a TD error. The TD error then reaches two sets of weights. Identify what each set of weights represents and what changes when it is updated.
Hints
- One set belongs to the critic and determines the value function.
- The other set belongs to the actor and determines the policy.
- State the different meaning of changing each set.
What do you think happens?
If the TD error is removed from the diagram after the critic produces it, which learning path is missing?
Reveal answer
Answer: Both the critic's and actor's updates
The TD error is the shared reinforcement signal. It changes the critic's value parameters and the actor's policy parameters.
Learning Rule Summary
- The critic evaluates states, produces state values, and creates the TD error.
- The actor uses state features to produce a k-dimensional action vector through its policy parameters.
- The same state features are supplied to both components through separate weighted connections.
- The critic combines reward with information about changes in estimated state values to produce the TD error.
- The TD error updates critic value parameters and actor policy parameters, giving the two updates different meanings.
Key Takeaways
- The critic evaluates states and produces the TD error.
- The actor produces the action vector and adjusts its policy.
- State features reach both components through weighted connections.
- Reward contributes to the critic's TD-error calculation rather than reaching the actor through a direct path.
- The shared TD error updates value parameters in the critic and policy parameters in the actor.