State-Visit Distributions
The policy gradient theorem links performance changes to the policy-weight vector θ.
Following Performance Changes
The policy gradient theorem is about a specific question: how does the performance of a policy change when its policy weights change? The weights are collected in a vector called θ. The theorem gives an analytic expression for the gradient of performance with respect to the components of θ.
The gradient described by the theorem is a column vector of partial derivatives with respect to the components of θ. It is not introduced as a gradient of the state-visit distribution itself.
Tracing Visits Through Episodes
To understand dπ(s), begin with a fixed starting state s0. Follow the policy through randomly generated episodes. In each episode, choose one particular state s and count the number of time steps during which the process is in that state. The value dπ(s) is the expected value of that count across the episodes.
What Determines dπ(s)
The state-visit distribution is produced by two interacting sources. The policy π determines behavior by choosing actions. The MDP dynamics determine how the episode develops after those actions. Starting from s0, the policy and the dynamics jointly generate the episodes from which state visits are counted.
This means dπ(s) cannot be understood by looking at the policy alone. The same action choices can lead to different episode developments under different MDP dynamics. Conversely, the same dynamics can produce different state-visit patterns when the policy changes.
A Worked Visit Count
Imagine several episodes that all start at s0. Choose a target state called s. In one generated episode, the process visits s twice. In another, it visits s once. In a third, it does not visit s. These are counts for individual episodes, not yet the value dπ(s). The expected value of the counts across the randomly generated episodes is the state-visit quantity dπ(s).
From Episode Counts to dπ(s)
Three episodes starting at s0 visit target state s two times, one time, and zero times respectively. What does dπ(s) represent?
Record each episode's count: The visit counts are two, one, and zero. Each count records how many time steps the process spent in state s during that episode.
Keep the episode-generating process fixed: The episodes are generated by following policy π together with the MDP dynamics from starting state s0.
Take the expected count: Across the randomly generated episodes, the expected value of the visit count is the quantity denoted dπ(s).
dπ(s) is the expected number of time-step visits to state s across episodes that start at s0 and follow both π and the MDP dynamics.
The example does not turn dπ(s) into a simple episode frequency. It counts time-step visits within episodes, then takes the expected value of those counts.
θ Versus dπ(s)
The vector θ controls the policy's weights. Changing θ can change the policy, and changing the policy can change policy performance. Because the policy also influences which episodes are generated, a change in θ can affect the state-visit distribution indirectly through changed behavior.
| Question | Policy gradient theorem | State-visit term |
|---|---|---|
| What is being described? | How policy performance changes with respect to θ | How often state s is visited on average |
| Main object | The performance gradient, a column vector | dπ(s), an expected time-step visit count |
| What generates the relevant episodes? | The policy and MDP dynamics | The policy and MDP dynamics |
| Is the theorem's key contribution a derivative of dπ(s)? | No; its expression does not require taking the derivative of the state distribution | dπ(s) is used as a state-visit quantity |
Common Interpretation Errors
Treating dπ(s) as an arbitrary list of state probabilities.
The source defines dπ(s) in the episodic setting as the expected number of time-step visits to state s.
Fix:
Think of dπ(s) as an expected visit count produced by episodes starting at s0 and following π and the MDP dynamics.Attributing the state distribution only to the policy.
The MDP dynamics also determine how the episode develops.
Fix:
Treat policy behavior and MDP dynamics as joint determinants of the episodes used to define dπ(s).Saying that the policy gradient theorem differentiates dπ(s) directly.
The theorem describes performance changes with respect to θ and does not require taking the derivative of the state distribution.
Fix:
Keep the target of differentiation clear: the theorem gives the gradient of performance with respect to the policy-weight vector θ.Forgetting that a state may be visited at multiple time steps in one episode.
The definition counts how many time steps are spent in state s.
Fix:
Count each relevant time-step visit before taking the expected value across episodes.
Check Your Understanding
Explain, in your own words, why dπ(s) depends on both π and the MDP dynamics. Then state what quantity the policy gradient theorem differentiates with respect to θ.
Hints
- Start with the fixed starting state s0.
- Ask who chooses actions and what determines how the episode develops.
- Distinguish the expected visit count from the performance gradient.
What do you think happens?
Suppose the policy changes because θ changes. Does that necessarily mean the policy gradient theorem must take the derivative of dπ(s) itself?
Reveal answer
Answer: No, the theorem gives performance change with respect to θ without requiring the derivative of the state distribution.
The state-visit distribution is determined by the policy and MDP dynamics, but the theorem's important contribution is that its performance-gradient expression does not require taking the derivative of that distribution.
Key Takeaways
- The policy gradient theorem describes how policy performance changes with respect to the policy-weight vector θ.
- The gradient is a column vector containing partial derivatives with respect to θ's components.
- dπ(s) is the expected number of time-step visits to state s across episodes starting at s0.
- The policy π and the MDP dynamics jointly determine the episodes from which dπ(s) is defined.
- The theorem uses the state-visit distribution without requiring a derivative of the state distribution itself.
Key Takeaways
- The policy gradient theorem concerns performance changes with respect to θ, not a direct derivative of dπ(s).
- dπ(s) summarizes the expected number of time-step visits to state s across episodes beginning at s0.
- Both policy action choices and MDP dynamics determine the episodes and therefore the state-visit distribution.
- A change in θ can change the policy and indirectly change visitation, while the theorem still avoids requiring a derivative of the state distribution.