Concepts / Value Functions and Policy Evaluation

Value Functions and Policy Evaluation

The actor changes the policy; the critic learns a value function for current behavior.

  • Programming

The Delayed-Reward Problem

In reinforcement learning, the useful consequence of a sequence of decisions may arrive considerably later than the decisions themselves. This creates a credit-assignment problem: the system must improve its current behavior even though the primary reward signal may not yet explain which earlier choices were useful. Actor-critic architecture addresses this timing problem by dividing the work between two components. The actor changes behavior, while the critic evaluates behavior as learning proceeds.

critic evaluatesguides policy adjustmenttask outcomePolicy decisionActor changes behaviorTD errorSecondary reward signalPrimary rewardMay arrive later
How can intermediate evaluations guide the actor before the environment provides the final primary reward?

Two Jobs in One Architecture

The actor and critic have separate roles. The actor changes the policy, meaning it changes the behavior the system follows. The critic learns a value function for the current behavior. Its value predictions evaluate the return associated with that current policy. Keeping these jobs distinct lets one component produce behavior while the other supplies an ongoing evaluation of that behavior.

changesis evaluated byproducesguidesActorChanges the policyCurrent behaviorPolicy being evaluatedCriticLearns a value functionTD errorSecondary reward signal
How do the actor's policy changes and the critic's value estimates interact during learning?
ComponentMain responsibilityLearning connection
ActorChanges the policyUses the critic's secondary reward signal to adjust behavior
CriticLearns a value function for current behaviorUses temporal-difference learning to evaluate the current policy

The actor changes behavior; the critic evaluates the behavior.

Following the Critic's Evaluation Loop

The critic uses temporal-difference learning as its evaluation mechanism. As the system follows its current policy, the critic maintains value predictions about the return associated with that behavior. A temporal-difference update uses information available during learning, including the current reward and the value estimate associated with the next state, to revise the critic's evaluation. The difference recorded by this revision is the temporal-difference error.

predictioncontributescontributesupdatessupports next evaluationCurrent stateCurrent value predictionCurrent rewardNew informationNext stateNext value estimateTD errorDifference in evaluationUpdated valuefunctionRevised critic
How does the critic use the current reward and the next state's value estimate to update its value function?

TD learning lets the critic evaluate the current policy through its value predictions. The critic is therefore not waiting only for the final primary reward; it is learning from changes in its evaluation as behavior unfolds.

From TD Error to Policy Feedback

The temporal-difference error does more than update the critic. It also supplies the actor with a secondary reward signal. This signal is an immediate evaluation derived from the critic's TD calculation. The actor can use it to adjust the policy before the task's primary reward necessarily arrives.

is evaluatedproducesguideschangesCurrent behaviorProduced by policyCriticEvaluates behaviorTD errorSecondary reward signalActorUpdates policy
How does the TD error move from the critic to guide the actor's policy update?

A Delayed Sequence in Practice

Updating Behavior Before the Final Outcome

Imagine a system following its current policy through a sequence of decisions. The primary reward for the whole sequence is delayed, so it is not immediately available to explain whether the current behavior was useful.

The actor acts: The actor follows the current policy and produces behavior through the sequence.

The critic evaluates: The critic applies temporal-difference learning to its value function while the sequence unfolds.

The prediction changes: New information changes the critic's evaluation of the return associated with the current policy.

The TD error is passed on: The difference in the critic's evaluation becomes a secondary reward signal for the actor.

The actor adjusts: The actor can update the policy using this immediate evaluation instead of waiting for the delayed primary reward alone.

The critic provides an intermediate teaching signal, allowing policy adjustment during a delayed-reward sequence.

start frompredictsformsStatePoint in behaviorCurrent policyBehavior to followExpected returnValue predictionPolicy evaluationCritic's assessment
How does a value function represent the expected future return of following the current policy from each state?

Hull's Secondary-Reinforcement Idea

The connection to Hull's hypothesis is theoretical. Hull's hypothesis concerns how animals learn when reinforcement is delayed. The actor-critic architecture provides a clear illustration of the proposed correspondence: the critic's value-function predictions change during learning, and the resulting TD error functions as secondary reinforcement for the actor. In this view, the secondary signal helps bridge the time between behavior and the later primary reward.

motivates ongoing evaluationchanges produceacts asillustratesDelayedreinforcementLearning challengeValue predictionCritic's evaluationSecondaryreinforcementSignal for policyadjustmentHull's hypothesisTheoretical correspondenceTD errorImmediate evaluation
How does the critic's TD error correspond to Hull's proposed secondary reinforcement signal?

The architecture is presented as a clear illustration of the theoretical correspondence, not as a claim that the TD error replaces the primary reward. Its role is to provide reinforcement for policy adjustment while the primary reward is delayed.

Mistakes About Actor-Critic Learning

  • Treating the actor and critic as if they perform the same job

    The actor changes the policy, while the critic learns a value function for current behavior.

    Fix: Keep the roles separate: the actor changes behavior and the critic evaluates that behavior.

  • Assuming the actor must wait for the final primary reward

    The critic's prediction changes provide a more immediate teaching signal through the TD error.

    Fix: Recognize the TD error as a secondary reward signal that can guide the actor before the primary reward necessarily arrives.

  • Calling the TD error the task's primary reward

    The TD error is derived from the critic's calculation and is described as a secondary signal.

    Fix: Describe it as an immediate evaluation that supplements, rather than replaces, the primary reward.

  • Describing TD learning only as an actor-update method

    TD learning is the evaluation mechanism inside the architecture and lets the critic evaluate the current policy.

    Fix: Start with the critic's value predictions, then explain how the resulting TD error guides the actor.

Check Your Understanding

MEDIUM

A task gives its primary reward only after a long sequence of decisions. Explain, in order, how the actor, the critic, TD learning, and the TD error work together so that the actor can receive useful feedback before the primary reward arrives.

Hints
  • Begin with the separate jobs of the actor and critic.
  • Identify what the critic learns and which algorithm supports that learning.
  • Explain why the TD error is called a secondary reward signal.
  • End by stating why this helps with delayed reinforcement.

What do you think happens?

If the primary reward is delayed, does actor-critic learning leave the actor without any evaluation until that reward arrives?

  • Yes, the actor must always wait
  • No, the critic can provide a secondary signal through its TD error
Reveal answer

Answer: No, the critic can provide a secondary signal through its TD error.

TD learning allows the critic's changing value predictions to produce a TD error, which supplies the actor with an immediate evaluation before the primary reward necessarily arrives.

Key Takeaways

  1. The actor changes the policy, while the critic learns a value function for the current behavior.
  2. TD learning is the critic's evaluation mechanism.
  3. The critic's TD error records a difference in evaluation and supplies the actor with a secondary reward signal.
  4. This secondary signal gives the actor feedback before a delayed primary reward necessarily arrives.
  5. Actor-critic architecture provides a clear illustration of the correspondence with Hull's hypothesis about learning under delayed reinforcement.

Key Takeaways

  • The actor changes behavior; the critic evaluates the current policy.
  • The critic uses TD learning to update its value function through ongoing evaluation.
  • The TD error acts as a secondary reward signal for the actor.
  • Secondary evaluation helps address credit assignment when the primary reward is delayed.
  • This actor-critic mechanism corresponds theoretically to Hull's hypothesis about delayed reinforcement.