Actor-Critic Algorithms
Non-contingent eligibility traces do not correlate actions with consequent changes in the reward signal.
Why Prediction Is Not Control
Actor-critic algorithms divide reinforcement learning into two connected responsibilities. The critic learns to predict the expected return, while the actor learns how actions should be selected. This distinction matters because a method can support prediction without providing enough information for control. In particular, non-contingent eligibility traces can be adequate for the critic's prediction task while failing to connect an individual action with a later change in the reward signal.
The Missing Action-to-Outcome Link
Imagine the learning process in two moments. First, an agent performs an action. Later, the common reinforcement signal changes. For control learning, the method must connect those two events so that the earlier action can be evaluated in light of its later consequence. Non-contingent eligibility traces do not provide this action-to-consequent-reward-change connection.
Checking whether control information exists
An agent performs an action, and the common reinforcement signal changes afterward. What must the learning process establish before that action can support control learning?
Identify the earlier event: The earlier event is the particular action performed by the agent.
Identify the later event: The later event is the change in the common reinforcement signal.
Check the connection: The learning process must relate the action to the later reward change. Non-contingent eligibility traces do not provide this connection.
Determine the consequence: Without that connection, the signal can support prediction for the critic but is not sufficient support for the actor's control learning.
Control requires an action-to-consequent-reward-change connection; non-contingent eligibility traces do not supply it.
Actor and Critic Responsibilities
The actor and critic are separate components with connected responsibilities. The actor follows a policy and makes action choices. It learns policies, so its role is control. The critic studies the policy currently being followed by the actor. It learns to predict the expected return and evaluates that current policy, so its role is prediction and evaluation.
| Component | Main responsibility | What it learns or provides |
|---|---|---|
| Critic | Prediction and evaluation | Predicts the expected return and evaluates the actor's current policy |
| Actor | Control | Learns policies and selects actions |
The two connected responsibilities in an actor-critic algorithm.
Following One Learning Cycle
A useful way to understand the algorithm is to trace one action choice. First, the actor follows its current policy and chooses an action. The critic examines the policy currently being followed and learns a state-value function for that policy using a temporal-difference method. The critic then sends a TD error, written as δ in the source material, to the actor. Finally, the actor uses that critique while continually updating its policy.
- The actor follows its current policy and makes an action choice.
- The critic evaluates the policy currently being followed by the actor.
- The critic uses a temporal-difference method to learn about that policy.
- The critic sends a TD error to the actor.
- The actor uses the TD error while continually updating its policy.
Reading the TD Error
The critic communicates its judgment through a TD error. At the beginner level, the most important feature is its sign. A positive TD error means that the action led to a state with a better-than-expected value. A negative TD error means that the action led to a state with a worse-than-expected value. The actor receives this feedback and uses it in its continuing policy updates.
Interpreting two TD-error signs
The actor chooses an action. In one case, the resulting state has a better value than the critic expected. In another case, the resulting state has a worse value than expected. Interpret the two TD errors.
Better result: When the resulting state has a better-than-expected value, the TD error is positive.
Worse result: When the resulting state has a worse-than-expected value, the TD error is negative.
Actor feedback: The sign communicates the critic's judgment to the actor, which uses the feedback while continually updating its policy.
Positive means better than expected; negative means worse than expected.
What do you think happens?
The critic observes that an action led to a state with a better-than-expected value. What sign should the TD error have?
Reveal answer
Answer: Positive
The source defines a positive TD error as indicating a better-than-expected result.
Team Learning and Differentiated Behavior
The defined team learning setting contains a population of agents that receives a common reinforcement signal. The goal is for the agents to learn differentiated patterns of activity. This is more demanding than making every agent predict the same quantity. The learning process must provide information about how each agent's activity relates to changes in the shared reward signal.
Shared prediction does not necessarily produce differentiated behavior. If agents receive or learn the same shared prediction but the learning process does not correlate each agent's actions with later changes in the common reward, the signal does not identify how individual activity should contribute to distinct patterns. This is why non-contingent eligibility traces are inadequate for the team setting: they do not supply the required action-to-reward-change connection for each agent.
Mistakes in Role Assignment
Treating prediction and control as the same job.
The source distinguishes the critic's prediction role from the actor's control role. Prediction alone does not provide the action-selection responsibility.
Fix:
Assign prediction and evaluation to the critic, and policy learning and action selection to the actor.Assuming that a common reinforcement signal identifies each agent's contribution.
The team setting requires a connection between each agent's actions and later changes in the shared reward signal.
Fix:
Check whether the learning method relates individual actions to their consequent reward changes.Reversing the meaning of the TD-error sign.
The source states that a positive TD error indicates a better-than-expected result, while a negative TD error indicates a worse-than-expected result.
Fix:
Remember: positive means better than expected; negative means worse than expected.Describing the critic as the component that directly controls the policy.
The actor follows a policy and makes action choices. The critic evaluates the policy currently followed by the actor.
Fix:
Describe the critic as evaluator and the actor as policy learner and controller.
Check Your Understanding
A population of agents receives one common reinforcement signal. Explain why a shared prediction alone may fail to produce differentiated patterns of activity. Then trace the actor-critic cycle: identify the actor's action, the critic's evaluation, the meaning of the TD error, and the actor's next learning responsibility.
Hints
- Start with the missing connection between an individual action and a later reward change.
- State separately what the critic predicts and what the actor controls.
- Use the sign of the TD error to describe whether the result was better or worse than expected.
- End by explaining that the actor uses the critic's feedback to continually update its policy.
- Name the component that selects an action according to a policy.
- Name the component that evaluates the policy currently being followed.
- Explain why non-contingent eligibility traces can support prediction but not sufficient control learning.
- Interpret a positive TD error.
- Interpret a negative TD error.
- Explain how the actor uses the TD error in its continuing policy updates.
Essential Takeaways
- Non-contingent eligibility traces do not correlate an action with a consequent change in the reward signal.
- The critic predicts the expected return and evaluates the policy currently followed by the actor.
- The actor controls action selection by learning and continually updating policies.
- A positive TD error indicates a better-than-expected result, while a negative TD error indicates a worse-than-expected result.
- In team learning, shared prediction alone does not guarantee differentiated agent behavior; individual action-to-reward-change information is required for control.
Key Takeaways
- Actor-critic algorithms separate prediction and control into a critic and an actor.
- The critic evaluates the actor's current policy using a temporal-difference method.
- The TD error tells the actor whether an action produced a better-than-expected or worse-than-expected result.
- Non-contingent eligibility traces fail to connect individual actions with later reward changes.
- Without that connection, a common reinforcement signal may support shared prediction without producing differentiated agent behavior.