Value Functions and Policy Evaluation
The actor changes the policy; the critic learns a value function for current behavior.
The Delayed-Reward Problem
In reinforcement learning, the useful consequence of a sequence of decisions may arrive considerably later than the decisions themselves. This creates a credit-assignment problem: the system must improve its current behavior even though the primary reward signal may not yet explain which earlier choices were useful. Actor-critic architecture addresses this timing problem by dividing the work between two components. The actor changes behavior, while the critic evaluates behavior as learning proceeds.
Two Jobs in One Architecture
The actor and critic have separate roles. The actor changes the policy, meaning it changes the behavior the system follows. The critic learns a value function for the current behavior. Its value predictions evaluate the return associated with that current policy. Keeping these jobs distinct lets one component produce behavior while the other supplies an ongoing evaluation of that behavior.
| Component | Main responsibility | Learning connection |
|---|---|---|
| Actor | Changes the policy | Uses the critic's secondary reward signal to adjust behavior |
| Critic | Learns a value function for current behavior | Uses temporal-difference learning to evaluate the current policy |
The actor changes behavior; the critic evaluates the behavior.
Following the Critic's Evaluation Loop
The critic uses temporal-difference learning as its evaluation mechanism. As the system follows its current policy, the critic maintains value predictions about the return associated with that behavior. A temporal-difference update uses information available during learning, including the current reward and the value estimate associated with the next state, to revise the critic's evaluation. The difference recorded by this revision is the temporal-difference error.
TD learning lets the critic evaluate the current policy through its value predictions. The critic is therefore not waiting only for the final primary reward; it is learning from changes in its evaluation as behavior unfolds.
From TD Error to Policy Feedback
The temporal-difference error does more than update the critic. It also supplies the actor with a secondary reward signal. This signal is an immediate evaluation derived from the critic's TD calculation. The actor can use it to adjust the policy before the task's primary reward necessarily arrives.
A Delayed Sequence in Practice
Updating Behavior Before the Final Outcome
Imagine a system following its current policy through a sequence of decisions. The primary reward for the whole sequence is delayed, so it is not immediately available to explain whether the current behavior was useful.
The actor acts: The actor follows the current policy and produces behavior through the sequence.
The critic evaluates: The critic applies temporal-difference learning to its value function while the sequence unfolds.
The prediction changes: New information changes the critic's evaluation of the return associated with the current policy.
The TD error is passed on: The difference in the critic's evaluation becomes a secondary reward signal for the actor.
The actor adjusts: The actor can update the policy using this immediate evaluation instead of waiting for the delayed primary reward alone.
The critic provides an intermediate teaching signal, allowing policy adjustment during a delayed-reward sequence.
Hull's Secondary-Reinforcement Idea
The connection to Hull's hypothesis is theoretical. Hull's hypothesis concerns how animals learn when reinforcement is delayed. The actor-critic architecture provides a clear illustration of the proposed correspondence: the critic's value-function predictions change during learning, and the resulting TD error functions as secondary reinforcement for the actor. In this view, the secondary signal helps bridge the time between behavior and the later primary reward.
The architecture is presented as a clear illustration of the theoretical correspondence, not as a claim that the TD error replaces the primary reward. Its role is to provide reinforcement for policy adjustment while the primary reward is delayed.
Mistakes About Actor-Critic Learning
Treating the actor and critic as if they perform the same job
The actor changes the policy, while the critic learns a value function for current behavior.
Fix:
Keep the roles separate: the actor changes behavior and the critic evaluates that behavior.Assuming the actor must wait for the final primary reward
The critic's prediction changes provide a more immediate teaching signal through the TD error.
Fix:
Recognize the TD error as a secondary reward signal that can guide the actor before the primary reward necessarily arrives.Calling the TD error the task's primary reward
The TD error is derived from the critic's calculation and is described as a secondary signal.
Fix:
Describe it as an immediate evaluation that supplements, rather than replaces, the primary reward.Describing TD learning only as an actor-update method
TD learning is the evaluation mechanism inside the architecture and lets the critic evaluate the current policy.
Fix:
Start with the critic's value predictions, then explain how the resulting TD error guides the actor.
Check Your Understanding
A task gives its primary reward only after a long sequence of decisions. Explain, in order, how the actor, the critic, TD learning, and the TD error work together so that the actor can receive useful feedback before the primary reward arrives.
Hints
- Begin with the separate jobs of the actor and critic.
- Identify what the critic learns and which algorithm supports that learning.
- Explain why the TD error is called a secondary reward signal.
- End by stating why this helps with delayed reinforcement.
What do you think happens?
If the primary reward is delayed, does actor-critic learning leave the actor without any evaluation until that reward arrives?
Reveal answer
Answer: No, the critic can provide a secondary signal through its TD error.
TD learning allows the critic's changing value predictions to produce a TD error, which supplies the actor with an immediate evaluation before the primary reward necessarily arrives.
Key Takeaways
- The actor changes the policy, while the critic learns a value function for the current behavior.
- TD learning is the critic's evaluation mechanism.
- The critic's TD error records a difference in evaluation and supplies the actor with a secondary reward signal.
- This secondary signal gives the actor feedback before a delayed primary reward necessarily arrives.
- Actor-critic architecture provides a clear illustration of the correspondence with Hull's hypothesis about learning under delayed reinforcement.
Key Takeaways
- The actor changes behavior; the critic evaluates the current policy.
- The critic uses TD learning to update its value function through ongoing evaluation.
- The TD error acts as a secondary reward signal for the actor.
- Secondary evaluation helps address credit assignment when the primary reward is delayed.
- This actor-critic mechanism corresponds theoretically to Hull's hypothesis about delayed reinforcement.