Gradient Ascent for Policy Optimization
The policy gradient theorem links performance changes to the policy-weight vector θ.
Why the Theorem Matters
A policy is controlled by a vector of weights, written as θ. Changing these weights can change the policy, and changing the policy can change its performance. The policy gradient theorem describes how performance changes with respect to those policy weights. Its gradient is a column vector containing the partial derivatives with respect to the components of θ.
The central quantity is not a derivative with respect to an individual state or an arbitrary list of state probabilities. It is the change in policy performance described as a function of the policy-weight vector θ. The theorem connects that performance change to the direction in which the policy weights influence the policy.
Following Visits Through Episodes
To understand dπ(s), imagine starting every episode at the same starting state s0. The policy π chooses actions, while the MDP dynamics determine how each episode develops. Because both the policy and the dynamics are involved, the resulting episodes can vary. For one particular state s, count how many time steps each episode spends in that state. The quantity dπ(s) is the expected value of that count across the randomly generated episodes.
Counting Visits to One State
Consider three generated episodes that all begin at s0. In these illustrative episodes, count the time steps spent in state s.
Episode 1: The sequence visits s twice, so its count for state s is 2.
Episode 2: The sequence visits s once, so its count for state s is 1.
Episode 3: The sequence does not visit s, so its count for state s is 0.
Expected count: Across these three illustrative episodes, the average count is (2 + 1 + 0) divided by 3, which is 1.
In this illustration, dπ(s) is 1. The value represents an expected number of time-step visits, not the visit count of one particular episode.
Reading dπ(s)
dπ(s) is the expected number of time-step visits to state s across episodes that start at s0 and follow both the policy π and the MDP dynamics.
The subscript π indicates that the visit count depends on the policy behavior. However, the policy is not the only influence. The MDP dynamics also determine how an episode unfolds after actions are chosen. Therefore, dπ(s) summarizes the combined effect of policy behavior and MDP dynamics on how often state s is encountered on average.
Separating the Gradient Targets
The policy gradient theorem differentiates performance with respect to the policy-weight vector θ. The resulting gradient is a column vector whose entries are partial derivatives with respect to the individual components of θ.
The theorem also uses dπ(s), because the expected state visits help describe the episodes produced by the policy and the MDP. But the theorem's important contribution is that the expression does not require taking the derivative of the state distribution. In other words, dπ(s) appears as part of the theorem's performance-gradient expression; it is not introduced as a separate gradient target in the theorem's policy-weight derivative.
| Quantity | Role in the theorem |
|---|---|
| Performance | The quantity whose change is described |
| θ | The policy-weight vector with respect to which the gradient is taken |
| Components of θ | The variables associated with the partial derivatives in the gradient column vector |
| dπ(s) | The expected time-step count for visits to state s |
| Derivative of dπ(s) | Not required by the theorem's expression |
What θ Controls
Changing One Policy Weight
Suppose a policy is controlled by a vector θ with several components. Consider changing one component while leaving the other components unchanged.
Weight change: The selected component of θ changes. Because θ controls the policy, this change can change the policy's behavior.
Episode consequences: A changed policy can produce different action choices. Together with the MDP dynamics, those choices can change how episodes develop and how often states are visited.
Performance consequence: Because the policy may have changed, its performance may also change.
Gradient interpretation: The corresponding component of the policy gradient describes performance change with respect to that component of θ. All such partial derivatives form the gradient column vector.
θ is the coordinate system for the policy-weight gradient: its components are the variables whose influence on performance is described.
This does not mean that each component of θ directly equals a state visit count. The weights control the policy, the policy helps determine behavior, and the policy together with the MDP dynamics determines the episodes summarized by dπ(s). The gradient then describes performance changes with respect to the original policy-weight components.
Mistakes to Avoid
Treating dπ(s) as the count from one episode
dπ(s) is the expected count across randomly generated episodes that start at s0 and follow the policy and MDP dynamics.
Fix:
Count visits in each episode, then interpret dπ(s) as the expected value of those counts.Attributing dπ(s) only to the policy
The MDP dynamics also determine how the episode develops after actions are chosen.
Fix:
Remember that policy behavior and MDP dynamics jointly determine the episodes used to define dπ(s).Thinking the theorem takes a separate gradient with respect to dπ(s)
The theorem's important contribution is that its expression does not require taking the derivative of the state distribution.
Fix:
Identify θ as the variable of differentiation and dπ(s) as the expected state-visit quantity used in the expression.Describing the gradient as one number
The gradient is a column vector containing partial derivatives with respect to the components of θ.
Fix:
Connect each entry of the gradient to a component of the policy-weight vector.
When reading a policy-gradient expression, label each quantity by its role: performance is the quantity whose change is described, θ supplies the policy-weight variables, and dπ(s) summarizes expected state visits under the policy and MDP dynamics.
Practice Check
An episodic process starts every episode at s0. The policy π and the MDP dynamics generate several episodes. In one episode, state s is visited three times; in another, it is visited zero times. Explain what additional information is needed to determine the expected count dπ(s), and identify the quantity with respect to which the policy gradient theorem differentiates performance.
Hints
- dπ(s) is based on expected counts across the randomly generated episodes, not just one or two selected counts.
- The gradient is taken with respect to the components of the policy-weight vector θ.
Key Takeaways
- The policy gradient theorem describes how policy performance changes with respect to the policy-weight vector θ.
- The gradient is a column vector of partial derivatives with respect to the components of θ.
- dπ(s) is the expected number of time-step visits to state s across episodes beginning at s0.
- The policy and the MDP dynamics jointly determine the episodes from which dπ(s) is defined.
- The theorem uses the state distribution without requiring a separate derivative of that state distribution.
Key Takeaways
- The policy gradient theorem connects changes in performance to the policy-weight vector θ.
- Its gradient contains partial derivatives for the components of θ.
- dπ(s) records the expected time-step count for visits to state s across episodes starting at s0.
- Both policy behavior and MDP dynamics determine the episodes summarized by dπ(s).
- The theorem does not require taking the derivative of the state distribution as a separate gradient target.