Generalized Policy Iteration: A Unified Framework
GPI converges through the interaction of policy evaluation and policy improvement.
Two Processes, One Solution
Generalized Policy Iteration, or GPI, converges through the repeated interaction of two processes: policy evaluation and policy improvement. Evaluation makes the value function consistent with the current policy. Improvement makes the policy greedy with respect to the current value function. Neither process is sufficient by itself to explain GPI convergence. The important idea is that each process changes the conditions used by the other.
Tracing the Dependency
The dependency runs in both directions. Improvement responds to the value function: it asks whether the current policy makes the best choices according to the values currently available. Evaluation responds to the policy: it asks whether the current value function correctly represents that policy. Because each side uses the other side's current state, changing either side can expose a problem in the other.
A Conceptual GPI Trace
Track what happens when a policy is evaluated and then improved.
Start with a policy: Begin with a current policy, labeled P. At this point, the policy is the object that evaluation will analyze.
Evaluate the policy: Policy evaluation produces a value function, labeled V, that is consistent with P. This gives improvement a value function to use.
Improve the policy: Policy improvement examines V and makes the policy greedy with respect to that value function. The policy may therefore change from P to a new policy.
Reevaluate after the change: The previous value function was consistent with the earlier policy, not necessarily with the changed policy. Evaluation must therefore respond to the new policy.
Continue until the requirements agree: The cycle continues until the policy is greedy with respect to its own evaluation function and the value function is the evaluation of that policy.
GPI is a feedback process: evaluation supports improvement, and improvement determines which policy must be evaluated next.
Why the Stages Temporarily Disagree
The two processes can temporarily pull in opposing directions because changing one can disturb the condition established by the other. Suppose evaluation has made a value function consistent with a policy. Improvement may then use that value function to change the policy. Once the policy changes, the existing value function may no longer be correct for the new policy. Evaluation must respond again.
The reverse tension also matters. Evaluation may update the value function while holding the policy in view. After that update, the policy may no longer be greedy with respect to the updated value function. Improvement then has a reason to change the policy. This is not a failure of GPI; it is the interaction that drives the process toward a joint solution.
The Stabilization Condition
GPI stabilizes when the policy is greedy with respect to its own evaluation function. At that point, the value function is the evaluation of the policy, and the policy is appropriate for that value function. The two objects no longer create a reason for one another to change.
This condition is stronger than saying that a policy was improved once or that a value function was evaluated once. It links the policy and value function to each other. Evaluation confirms that the value function represents the policy, while improvement confirms that the policy is greedy with respect to that value function. Stabilization requires both relationships at the same time.
From Stability to Optimality
At stabilization, the Bellman optimality equation holds. In this framework, that equation is the mathematical condition confirming that the policy-value pair has reached the joint solution. The policy is greedy with respect to its own evaluation function, and the value function evaluates that policy. Therefore, the stabilized policy and value function are optimal.
The key connection is therefore a chain of ideas: evaluation makes the value function consistent with the policy; improvement makes the policy greedy with respect to the value function; stabilization means both statements hold together; and the Bellman optimality equation holds at that stabilized point.
Common Misreadings
Treating policy evaluation and policy improvement as independent processes
Evaluation uses the current policy, while improvement uses the current value function. Changing either object changes the conditions used by the other.
Fix:
Understand GPI as repeated interaction: evaluate the policy, improve the policy, and evaluate again when improvement changes it.Calling GPI converged after one policy improvement
The existing value function may no longer be correct for the changed policy.
Fix:
Check whether the policy is greedy with respect to its own evaluation function and whether the value function evaluates that policy.Calling GPI converged after one evaluation
The updated value function may show that the policy is not greedy with respect to it.
Fix:
Use improvement to test whether the policy remains appropriate for the updated value function.Interpreting temporary disagreement as failure
Temporary tension is expected because each process can disturb the condition established by the other.
Fix:
Treat the disagreement as part of the feedback process and continue until both requirements hold together.Using the Bellman optimality equation without explaining stabilization
The equation matters here because it holds when the policy and value function have reached their joint solution.
Fix:
Connect the equation to the condition that the policy is greedy with respect to its own evaluation function.
Practice Check
A value function has just been made consistent with the current policy. Explain why policy improvement may still change that policy, and state what must happen after the change before GPI can be considered stabilized.
Hints
- Ask what improvement checks against the current value function.
- Ask whether the old value function is automatically correct for a changed policy.
- State the joint condition involving greediness and evaluation.
What do you think happens?
If evaluation makes the value function consistent with the current policy, does that alone guarantee GPI stabilization?
Reveal answer
Answer: No, because improvement must also find the policy greedy with respect to that value function.
Evaluation establishes consistency with the current policy, but the updated value function may show that the policy is not greedy. Stabilization requires both evaluation and improvement requirements to hold together.
Key Takeaways
- GPI converges through the repeated interaction of policy evaluation and policy improvement.
- Evaluation makes the value function consistent with the current policy, while improvement makes the policy greedy with respect to the current value function.
- Changing one process can disturb the condition established by the other, so temporary disagreement is expected.
- Stabilization occurs when the policy is greedy with respect to its own evaluation function and the value function evaluates that policy.
- At stabilization, the Bellman optimality equation holds, confirming that the policy-value pair is optimal.
Key Takeaways
- GPI is a feedback process, not a one-way sequence in which one process permanently finishes before the other begins.
- Policy evaluation and policy improvement affect one another because evaluation depends on the policy and improvement depends on the value function.
- Temporary conflict disappears when the policy is greedy with respect to its own evaluation function.
- That joint stabilization is the condition under which the Bellman optimality equation holds and the policy-value pair is optimal.