Policy Gradient Theorem
A policy-weight change has both an action-selection consequence and a state-distribution consequence.
The Hidden Consequence of a Weight Change
A policy weight change does not affect only the action chosen at the current state. It can also influence which states the agent reaches later. This gives the policy gradient problem two connected consequences: an action-selection consequence and a state-distribution consequence.
Tracing the Two Consequences
A conceptual policy change
Suppose a policy weight changes. Trace the two ways this change can affect performance.
First consequence: The changed weight alters the policy's action selection. This is the direct response of the policy parameterization.
Second consequence: Because the selected action can change, the agent can reach different states later. The policy has therefore influenced the distribution of states in which future actions are selected.
Performance implication: Performance depends on both the actions selected and the states in which those actions are selected. Looking only at the immediate action response is therefore insufficient.
One policy-weight change can have a direct action-selection effect and an indirect state-distribution effect.
The first consequence is relatively straightforward to connect to the policy parameterization: the weights affect action selection. The second consequence follows a longer chain. Changed actions can influence later state visits, and those visits determine where future actions are selected.
Where the Environment Enters
The policy parameterization exposes how a weight change affects action selection. The environment governs the other part of the chain: how those altered choices affect the states the agent reaches. As a result, the state-distribution consequence is determined by the environment and is typically unknown.
This is why the state-distribution effect cannot usually be read directly from the policy weights. Knowing how the weights alter action selection does not by itself reveal how the environment will determine the later states. The environment makes that part of the chain difficult to determine.
Why the Gradient Is Difficult
Estimating the performance gradient with respect to policy weights requires more than calculating the immediate action response. The estimate must account for how the policy weights affect action selection and how those altered selections affect the distribution of states encountered. The first part can be computed relatively straightforwardly from the policy parameterization. The second part is typically unknown because it is determined by the environment.
Mistakes in Reading the Theorem
Treating a policy-weight change as an action-selection effect only.
The changed action can influence which states the agent reaches later, creating a state-distribution consequence.
Fix:
Trace both the immediate action response and the later effect on the states encountered.Assuming that the state-distribution consequence can be computed from the policy parameterization alone.
The state-distribution consequence is determined by the environment and is typically unknown.
Fix:
Separate what the policy parameterization exposes from what the environment governs.Estimating the performance gradient from the immediate action response alone.
Performance depends on both selected actions and the states in which those actions are selected.
Fix:
Include the unknown state-distribution effect when reasoning about the difficulty of gradient estimation.
Check Your Understanding
A policy weight changes and the policy begins selecting actions differently. Explain why this change may affect the states encountered later. Then identify which part of this chain is relatively straightforward to compute from the policy parameterization and which part is typically unknown.
Hints
- Start with the direct effect on action selection.
- Then follow how altered choices can influence later states.
- The environment determines the difficult, typically unknown part.
What do you think happens?
If you know exactly how a policy weight changes action selection, is that enough to determine the full performance gradient?
Reveal answer
Answer: No, because the weight can also change the distribution of states encountered.
The direct action-selection consequence is only one part of the performance change. The state-distribution consequence is determined by the environment and is typically unknown, which makes the full gradient difficult to estimate.
Key Takeaways
- Changing policy weights can alter both action selection and the distribution of states encountered.
- The action-selection consequence is relatively straightforward to compute from the policy parameterization.
- The state-distribution consequence is determined by the environment and is typically unknown.
- Because performance depends on both actions and the states in which they are selected, the performance gradient is difficult to estimate.
- Looking only at the immediate action response is insufficient for understanding the full effect of a policy-weight change.
Key Takeaways
- A policy-weight change has both an action-selection consequence and a state-distribution consequence.
- Changed actions can influence which states the agent reaches later.
- The policy parameterization exposes the action response, while the environment governs the state-distribution response.
- The unknown state-distribution consequence is the central reason that estimating the performance gradient is difficult.