Concepts / Generalized Policy Iteration: A Unified Framework

Generalized Policy Iteration: A Unified Framework

GPI converges through the interaction of policy evaluation and policy improvement.

  • Programming

Two Processes, One Solution

Generalized Policy Iteration, or GPI, converges through the repeated interaction of two processes: policy evaluation and policy improvement. Evaluation makes the value function consistent with the current policy. Improvement makes the policy greedy with respect to the current value function. Neither process is sufficient by itself to explain GPI convergence. The important idea is that each process changes the conditions used by the other.

current policyproducescurrent valuesupdatesCurrent policychoices to evaluatePolicy evaluationmake value consistentValue functioncurrent estimatesPolicy improvementmake choices greedy
How does updating the value function change the policy, and how does the improved policy change the next evaluation?

Tracing the Dependency

The dependency runs in both directions. Improvement responds to the value function: it asks whether the current policy makes the best choices according to the values currently available. Evaluation responds to the policy: it asks whether the current value function correctly represents that policy. Because each side uses the other side's current state, changing either side can expose a problem in the other.

A Conceptual GPI Trace

Track what happens when a policy is evaluated and then improved.

Start with a policy: Begin with a current policy, labeled P. At this point, the policy is the object that evaluation will analyze.

Evaluate the policy: Policy evaluation produces a value function, labeled V, that is consistent with P. This gives improvement a value function to use.

Improve the policy: Policy improvement examines V and makes the policy greedy with respect to that value function. The policy may therefore change from P to a new policy.

Reevaluate after the change: The previous value function was consistent with the earlier policy, not necessarily with the changed policy. Evaluation must therefore respond to the new policy.

Continue until the requirements agree: The cycle continues until the policy is greedy with respect to its own evaluation function and the value function is the evaluation of that policy.

GPI is a feedback process: evaluation supports improvement, and improvement determines which policy must be evaluated next.

current policyproducesguidesupdatesnew evaluationPolicycurrent choicesEvaluationvalue matches policyValue functionevaluation resultImprovementpolicy becomes greedyNext policyevaluate again if changed
What happens first, what happens next, and how do evaluation and improvement alternate until convergence?

Why the Stages Temporarily Disagree

The two processes can temporarily pull in opposing directions because changing one can disturb the condition established by the other. Suppose evaluation has made a value function consistent with a policy. Improvement may then use that value function to change the policy. Once the policy changes, the existing value function may no longer be correct for the new policy. Evaluation must respond again.

The reverse tension also matters. Evaluation may update the value function while holding the policy in view. After that update, the policy may no longer be greedy with respect to the updated value function. Improvement then has a reason to change the policy. This is not a failure of GPI; it is the interaction that drives the process toward a joint solution.

evaluationimprovementrequires reevaluationPolicy Pcurrent policyUpdated policychanged by improvementValue function Vconsistent with PPrevious Vmay not match updatedpolicy
Why can evaluation favor the current policy while improvement immediately changes that policy before they eventually agree?

The Stabilization Condition

GPI stabilizes when the policy is greedy with respect to its own evaluation function. At that point, the value function is the evaluation of the policy, and the policy is appropriate for that value function. The two objects no longer create a reason for one another to change.

This condition is stronger than saying that a policy was improved once or that a value function was evaluated once. It links the policy and value function to each other. Evaluation confirms that the value function represents the policy, while improvement confirms that the policy is greedy with respect to that value function. Stabilization requires both relationships at the same time.

correct value for policycheck policy against valueno further conflictif policy changesTemporary mismatchpolicy and values pullapartEvaluationvalue matches policyImprovementpolicy matches valuesStable pairboth requirements hold
What changes from one iteration to the next, and how can we see when both the policy and value function stop changing?

From Stability to Optimality

At stabilization, the Bellman optimality equation holds. In this framework, that equation is the mathematical condition confirming that the policy-value pair has reached the joint solution. The policy is greedy with respect to its own evaluation function, and the value function evaluates that policy. Therefore, the stabilized policy and value function are optimal.

The key connection is therefore a chain of ideas: evaluation makes the value function consistent with the policy; improvement makes the policy greedy with respect to the value function; stabilization means both statements hold together; and the Bellman optimality equation holds at that stabilized point.

greedy conditionevaluation conditionsatisfiesconfirmsPolicygreedy with respect tovalueValue functionevaluation of policyStabilized pairrequirements hold togetherBellman optimalityequationmathematical confirmationOptimal pairpolicy and value function
How does a stabilized policy-value pair satisfy the Bellman optimality relationship?

Common Misreadings

  • Treating policy evaluation and policy improvement as independent processes

    Evaluation uses the current policy, while improvement uses the current value function. Changing either object changes the conditions used by the other.

    Fix: Understand GPI as repeated interaction: evaluate the policy, improve the policy, and evaluate again when improvement changes it.

  • Calling GPI converged after one policy improvement

    The existing value function may no longer be correct for the changed policy.

    Fix: Check whether the policy is greedy with respect to its own evaluation function and whether the value function evaluates that policy.

  • Calling GPI converged after one evaluation

    The updated value function may show that the policy is not greedy with respect to it.

    Fix: Use improvement to test whether the policy remains appropriate for the updated value function.

  • Interpreting temporary disagreement as failure

    Temporary tension is expected because each process can disturb the condition established by the other.

    Fix: Treat the disagreement as part of the feedback process and continue until both requirements hold together.

  • Using the Bellman optimality equation without explaining stabilization

    The equation matters here because it holds when the policy and value function have reached their joint solution.

    Fix: Connect the equation to the condition that the policy is greedy with respect to its own evaluation function.

Practice Check

MEDIUM

A value function has just been made consistent with the current policy. Explain why policy improvement may still change that policy, and state what must happen after the change before GPI can be considered stabilized.

Hints
  • Ask what improvement checks against the current value function.
  • Ask whether the old value function is automatically correct for a changed policy.
  • State the joint condition involving greediness and evaluation.

What do you think happens?

If evaluation makes the value function consistent with the current policy, does that alone guarantee GPI stabilization?

  • Yes, because evaluation is the only process that changes the value function.
  • No, because improvement must also find the policy greedy with respect to that value function.
  • Yes, because a policy cannot change after evaluation.
  • No, because the value function is irrelevant after evaluation.
Reveal answer

Answer: No, because improvement must also find the policy greedy with respect to that value function.

Evaluation establishes consistency with the current policy, but the updated value function may show that the policy is not greedy. Stabilization requires both evaluation and improvement requirements to hold together.

Key Takeaways

  1. GPI converges through the repeated interaction of policy evaluation and policy improvement.
  2. Evaluation makes the value function consistent with the current policy, while improvement makes the policy greedy with respect to the current value function.
  3. Changing one process can disturb the condition established by the other, so temporary disagreement is expected.
  4. Stabilization occurs when the policy is greedy with respect to its own evaluation function and the value function evaluates that policy.
  5. At stabilization, the Bellman optimality equation holds, confirming that the policy-value pair is optimal.

Key Takeaways

  • GPI is a feedback process, not a one-way sequence in which one process permanently finishes before the other begins.
  • Policy evaluation and policy improvement affect one another because evaluation depends on the policy and improvement depends on the value function.
  • Temporary conflict disappears when the policy is greedy with respect to its own evaluation function.
  • That joint stabilization is the condition under which the Bellman optimality equation holds and the policy-value pair is optimal.