Policy Gradient
Policy-gradient actor-critic learning combines a policy-based actor with a value-based critic.
Following the Learning Signal
A policy-gradient method searches through policies defined by numerical parameters. Instead of treating a value estimate as the primary object being searched, it estimates a direction for changing the policy parameters so that the policy can perform better. The direction comes from the agent's interaction with its environment.
The central search object is the policy defined by θ. A value estimate may help estimate a better direction, but value functions are not required by every policy-gradient method.
The Actor-Critic Learning Loop
The policy-gradient actor-critic method combines two learning processes. The actor is a differentiable policy written as π(a|s,θ). It selects actions and changes the policy parameters θ. The critic is a differentiable state-value function written as v̂(s,w). It estimates the value of the current state and changes the value-function parameters w. The actor and critic therefore have separate parameter sets, traces, and step sizes.
- Begin with the current state S.
- Select an action according to the current policy, written as A ∼ π(·|S,θ).
- Execute the action and observe the next state S′ and reward R.
- Use the reward, the next-state value estimate, and the current-state value estimate to compute the temporal-difference error δ.
- Update the critic trace ew and the actor trace eθ.
- Use the separate traces to change w and θ.
One Transition Through the System
A single actor-critic transition
Trace what happens when the agent starts in state S, selects an action using its current policy, and then observes S′ and R.
Policy selection: The actor uses π(a|s,θ) to select an action from the current state. The selected action is generated by the policy currently defined by θ.
Environment response: The environment returns the next state S′ and reward R. These observations are the experience used by the learning process.
Shared error: The method compares the reward plus the discounted value estimate of S′ with the current state's value estimate. The resulting temporal-difference error is δ.
Separate traces: The critic updates ew using the current value-function gradient, while the actor updates eθ using the gradient of the log policy for the selected action. Each trace also retains a decayed contribution from earlier experience.
Separate parameter changes: The critic uses δ and ew to change w. The actor uses δ and eθ to change θ. The same δ reaches both branches, but the branches change different parameter sets.
One interaction produces a shared learning signal but two distinct parameter updates: one for the value estimate and one for the policy.
The Shared Role of δ
The temporal-difference error δ connects the actor and critic learning processes. It compares the reward plus the discounted value estimate of the next state with the current state's value estimate. Once the traces have been formed, δ scales their influence. The critic uses that scaled influence to change w, while the actor uses it to change θ.
Suppose a value estimate looks wrong. The first branch to inspect is the critic path involving v̂, ew, β, and w. If the policy parameters look wrong, inspect the actor path involving π, eθ, α, and θ. If both paths are affected, inspect δ because both updates share it.
Eligibility Traces and Recent History
Eligibility traces prevent the update from depending only on the immediately current state or action. Each trace combines a decayed version of its previous value with a current gradient contribution. The critic trace uses the current value-function gradient. The actor trace uses the gradient of the log policy for the selected action. The λ terms control how much of the previous trace remains, while the current gradients add the present state's contribution.
Why an earlier experience can still matter
Imagine that the agent visits one state, then moves to another state before an update is evaluated.
First visit: The first state's current gradient contributes to the relevant actor or critic trace.
Later visit: At the next step, the previous trace is not simply discarded. Its contribution is decayed according to λ, and the new state's gradient is added.
Current error: When δ is applied, the retained contribution from the earlier step can influence the current update along with the newer contribution.
Eligibility traces carry a decayed record of recent states or actions, allowing current updates to reflect more than the most recent experience.
Two Updates, Two Parameter Sets
| Learning branch | Parameters changed | Trace used | Step size | Gradient source |
|---|---|---|---|---|
| Actor | θ | eθ | α | Gradient of the log policy for the selected action |
| Critic | w | ew | β | Current value-function gradient |
What the Method Searches
A policy-gradient method searches over parameter-defined policies. The numerical parameters θ define the policy being followed. Experience from interaction with the environment provides information for estimating a direction in which those parameters should move. Repeating this adjustment produces a sequence of parameter-defined policies rather than a search whose primary object is a value estimate.
Interaction matters because the improvement direction is estimated from what happens when the current policy acts in the environment. Without the resulting states and rewards, the method would not have the experience used to form the temporal-difference error and update its parameters.
Mistakes in Tracing Updates
Treating θ and w as the same parameter set
The actor changes θ, while the critic changes w. They are separate learning processes.
Fix:
Track the actor and critic branches separately before applying the shared δ signal.Using the critic trace for the actor update
The actor uses eθ, while the critic uses ew.
Fix:
Match θ with eθ and α; match w with ew and β.Ignoring the environment interaction
The direction is estimated from the agent's interaction with the environment.
Fix:
Start the trace with the current state, selected action, next state, and reward.Assuming eligibility traces store only the current experience
Eligibility traces combine a decayed previous contribution with the current gradient.
Fix:
Follow the λ-controlled decay and then add the present state's or action's contribution.Assuming every policy-gradient method requires a value function
Value estimates can improve gradient estimates in some methods, but they are not required in all of them.
Fix:
Describe the value function as part of the actor-critic combination, not as a universal requirement.
Check the Mechanism
A learner says: “The actor and critic both use δ, so they must update the same parameters.” Explain why this is incorrect. Name the parameter set, trace, and step size used by each branch.
Hints
- Begin with the actor's parameter set and trace.
- Then identify the critic's parameter set and trace.
- Remember that δ is shared, but the destinations of the updates are different.
Trace a two-step interaction in words. Include the first state and action, the later state and reward, the decayed previous eligibility trace, the current gradient contribution, and the way δ affects the final actor and critic updates.
Hints
- Separate the actor trace from the critic trace.
- Use λ to explain what happens to the earlier contribution.
- End by naming which parameter set each branch changes.
Key Takeaways
- Policy-gradient methods search over policies defined by numerical parameters and estimate directions for changing those parameters.
- In an actor-critic method, the actor uses π(a|s,θ) and changes θ, while the critic uses v̂(s,w) and changes w.
- The temporal-difference error δ connects the two branches and scales both updates.
- Eligibility traces preserve decayed contributions from recent states or actions and combine them with current gradients.
- Value-function estimates can improve policy-gradient estimates in some methods, but they are not required by every policy-gradient method.
Key Takeaways
- An actor-critic policy-gradient method combines a policy-based actor with a value-based critic.
- The actor updates θ, the critic updates w, and each uses its own trace and step size.
- The temporal-difference error δ is shared by both learning branches.
- Eligibility traces carry decayed information from recent experiences into current updates.
- Policy-gradient methods search through parameter-defined policies using directions estimated from environment interaction.