Policy Weights
The theorem connects performance to policy weights through an analytic gradient expression.
The Quantity Gradient Ascent Needs
A policy controlled by adjustable weights needs a way to connect those weights to performance. The policy gradient theorem provides an analytic expression for the gradient of performance with respect to the policy weights. That gradient is the quantity gradient ascent needs to approximate.
Keep two ideas separate. The policy gradient theorem describes the gradient relationship. Gradient ascent is the optimization method that uses an estimate or approximation of that gradient to seek better performance. The theorem is therefore relevant to gradient ascent, but it is not itself the optimization method.
What the Theorem Leaves Out
The policy gradient theorem has an important omission: its expression does not involve the derivative of the state distribution. This absent derivative is part of what distinguishes the theorem from an explanation that incorrectly adds a derivative of the state distribution.
Following One REINFORCE Update
A REINFORCE update can be understood as combining three ingredients: the selected action's return, the eligibility vector, and the action probability. The return controls how strongly the selected action's update is applied. The eligibility vector identifies the direction that increases the probability of repeating that action in the state. The probability denominator balances the effect of action frequency.
A qualitative update trace
Suppose a policy selects an action in a state, and that action produces a return. Trace what determines the resulting policy-weight change.
Start with the selected action: The update concerns the action the policy selected in the state.
Use the return: The selected action's return determines the strength of the update. A return is therefore not merely a record of what happened; it influences how much the policy is changed.
Use the eligibility vector: The eligibility vector supplies the direction associated with increasing the probability of repeating the selected action in that state.
Account for action probability: The selected action's probability appears in the denominator, balancing the effect of action frequency rather than simply rewarding an action because it was selected often.
Combine the influences: The return affects update strength, the eligibility vector affects update direction, and the probability denominator helps balance the update.
The policy-weight change is determined by the return and the policy's sensitivity to the selected action, with the action probability providing the denominator that balances frequency.
Reading the Eligibility Vector
The eligibility vector is the gradient of the log probability of the selected action. It identifies the direction in which the policy weights should change to increase the probability of repeating that action in the state.
The eligibility vector is about sensitivity, not about how often an action has been selected. It describes how the selected action's log probability changes with the policy weights. That is why it supplies a direction for changing the weights rather than acting as a simple frequency counter.
Why Probability Is in the Denominator
The action probability appears in the denominator because differentiating the log probability produces the derivative of the action probability divided by that probability. In the REINFORCE interpretation, this denominator balances the effect of action frequency. The update therefore depends on more than how many times an action was selected.
Frequent Selection Is Not Enough
A common shortcut is to assume that REINFORCE simply strengthens whichever actions are selected most often. That interpretation is incorrect. The update depends on both the return produced by the action and how the policy probability changes with its parameters.
Assuming the most frequently selected action automatically receives the largest update.
REINFORCE does not simply reinforce selection frequency. The update also depends on the return and on how the policy probability changes with its parameters.
Fix:
Evaluate the selected action's return and eligibility vector, while accounting for the action probability in the denominator.Treating the policy gradient theorem as the optimization method.
The theorem provides the analytic gradient expression; gradient ascent is the method that needs an approximation of that gradient.
Fix:
Describe the theorem as the source of the gradient relationship and gradient ascent as the optimization method that uses it.Adding a derivative of the state distribution to the policy gradient theorem.
The derivative of the state distribution does not appear in the policy gradient theorem described here.
Fix:
Remember the theorem's stated performance-with-respect-to-policy-weights gradient and its omission of the state-distribution derivative.
Check Your Understanding
A learner says: The policy selected action A many times, so REINFORCE must increase its weight substantially. Correct the statement using the roles of return, eligibility vector, and action probability.
Hints
- Ask what determines the strength of the update.
- Ask what identifies the update direction.
- Explain why frequency alone is not the complete rule.
State the derivative that does not appear in the policy gradient theorem, then explain why confusing its absence matters.
Hints
- The missing quantity concerns the state distribution.
- The omission helps distinguish the theorem from an incorrect expanded explanation.
Key Takeaways
- The policy gradient theorem gives an analytic expression for the gradient of performance with respect to policy weights.
- That gradient is the quantity gradient ascent needs to approximate; the theorem and the optimization method are distinct.
- The derivative of the state distribution does not appear in the policy gradient theorem.
- In REINFORCE, the return influences update strength and the eligibility vector, defined as the gradient of the selected action's log probability, supplies update direction.
- REINFORCE does not simply strengthen the most frequently selected action; return, policy sensitivity, and the action-probability denominator all matter.
Key Takeaways
- The policy gradient theorem connects performance to adjustable policy weights through an analytic gradient expression.
- Gradient ascent needs an approximation of that gradient, but the theorem is not itself the optimization method.
- The theorem does not include the derivative of the state distribution.
- A REINFORCE update uses the return, the eligibility vector, and the selected action's probability.
- Frequently selected actions are not automatically reinforced more strongly.