Concepts / Generalized Policy Iteration

Generalized Policy Iteration

Evaluation and improvement impose different requirements on a policy-value pair.

  • Programming

Two Processes, One Search

Generalized Policy Iteration, or GPI, is a search conducted by two interacting processes. Policy evaluation asks the value function to match the current policy. Policy improvement asks the policy to become greedy with respect to the current value function. Because each process changes one side of the policy-value pair, progress by one process can temporarily disturb the condition that the other process was establishing.

The apparent conflict is temporary. The long-run purpose of both processes is to reach one joint solution: an optimal policy together with an optimal value function.

Tracing the First Interaction

improvement reads Vevaluation responds to Pupdated values affect greedinessPolicy Pcurrent choicesImproved Pgreedy with respect to VValue Vcurrent estimatesEvaluated Vconsistent with improved P
How do evaluation and improvement change the policy-value pair in alternating steps?

An Abstract P-V Trace

Trace what happens when improvement changes policy P and evaluation then responds.

Start: Begin with a policy P and a value function V. At this point, V represents the current policy to the extent established by evaluation.

Improve: Improvement examines the current value function and changes P so that the policy becomes greedy with respect to that value function.

Disturb: The changed policy is not necessarily the policy that the existing V evaluated. The previous V can therefore become inaccurate for the changed policy.

Evaluate: Evaluation responds to the changed P by making the value function consistent with that policy.

Continue: The updated value function can affect whether P is still greedy. The two processes continue interacting until their requirements are satisfied together.

Improvement responds to the value function, while evaluation responds to the policy. Neither process can be understood as permanently independent of the other.

The important direction of dependence is simple: improvement reads the value function and changes the policy; evaluation reads the changed policy and adjusts the value function. A policy change can invalidate the value function's fit to that policy, so evaluation is needed again.

Two Requirements in Tension

updated V can affect greedinesschanged P can affect consistencyrepeated interactionEvaluationV matches PTemporarydisagreementone change can disturb theotherJoint solutionoptimal P and VImprovementP is greedy for V
What two requirements must a policy-value pair satisfy, and why can they temporarily disagree?

Evaluation and improvement impose different requirements on the same pair of objects. Evaluation requires the value function to be consistent with the current policy. Improvement requires the policy to be greedy with respect to the current value function. These requirements can temporarily pull in opposing directions because satisfying one can disturb the other.

When One Side Disturbs the Other

V is consistent with Pfit may be inaccuratePolicy Pold policyPolicy P'changed policyValue Vevaluates PValue Vold fit
What changes when improvement changes the policy while the value function still describes the old policy?

Suppose evaluation has made V consistent with policy P. Improvement then uses V to change P. The existing V has not automatically been recomputed for the changed policy. It may therefore describe the old policy rather than the new one. This is the first form of temporary conflict: improvement makes progress toward greediness, but its policy change can make the current value function inaccurate.

evaluation responds to Pimprovement responds to VEvaluationchanges VCurrent policyV matches PImprovementchanges PCurrent valueP is greedy for V
How do the two processes respond to different sides of the policy-value pair?

The reverse disturbance is also possible. Evaluation can make V consistent with the current policy, but the updated V can change the policy's relationship to the values. A policy that was greedy with respect to an earlier value function may no longer be greedy with respect to the newly evaluated one. Thus evaluation can remove the condition that improvement was trying to establish.

The Stabilized Pair

evaluationimprovementno further disturbancePolicy Pcurrent policyGreedy Pwith respect to VStable pairoptimal policy and valueValue Vevaluation of P
What does it look like when neither process produces a further change?

GPI stabilizes when the policy is greedy with respect to its own evaluation function. This condition links the two objects rather than declaring one process complete in isolation: the value function is the evaluation of the policy, and the policy is greedy with respect to that value function.

From Stability to Optimality

both requirements concern the pairstabilizationconfirms optimalityV evaluates Pevaluation requirementP is greedy for Vimprovement requirementBellman optimalityequationequation (4.1)Optimal P and Vjoint solution
How does the stabilized policy-value pair connect to the Bellman optimality equation?

At stabilization, the Bellman optimality equation, identified in the source as equation (4.1), holds. The equation is the mathematical confirmation of the joint solution: the stabilized value function and policy are optimal. The interaction between evaluation and improvement explains how the pair reaches this condition; the equation identifies what is true once the condition has been reached.

Common Reasoning Errors

  • Treating policy improvement as a permanent final step.

    Changing the policy can make the current value function inaccurate for that changed policy.

    Fix: Recognize that evaluation must respond to the changed policy.

  • Assuming evaluation can only help improvement.

    The updated value function can cause the policy to stop being greedy with respect to it.

    Fix: Check the relationship between the policy and its updated evaluation.

  • Describing GPI as two independent procedures.

    Each process changes the conditions used by the other.

    Fix: Track the feedback: improvement responds to V, and evaluation responds to P.

  • Confusing temporary disagreement with the final goal.

    The processes cooperate toward an optimal value function and an optimal policy.

    Fix: Separate local disruption from the eventual joint solution.

  • Defining convergence as one process completing first.

    Convergence requires the policy to be greedy with respect to its own evaluation function.

    Fix: Look for simultaneous satisfaction of both requirements.

Check Your Reasoning

MEDIUM

A value function V is consistent with policy P. Improvement changes P into a new policy. Explain why V may now be inaccurate, what evaluation must do next, and why the newly evaluated V might require another improvement check.

Hints
  • Start with the fact that V represented the earlier policy.
  • Identify which process responds to the changed policy.
  • Finish by checking whether the policy is greedy with respect to the updated value function.

What do you think happens?

If evaluation has made V consistent with P, does that alone guarantee that P is greedy with respect to the updated V?

  • Yes, evaluation and improvement impose the same requirement
  • No, evaluation can change the value function and remove the policy's greediness
  • Only if the policy was changed before evaluation
Reveal answer

Answer: No, evaluation can change the value function and remove the policy's greediness.

Evaluation makes the value function consistent with the policy. Improvement separately asks whether the policy is greedy with respect to that value function. Because the updated value function can alter that relationship, both requirements must be checked together.

Key Takeaways

  1. Evaluation makes the value function consistent with the current policy.
  2. Improvement makes the policy greedy with respect to the current value function.
  3. Changing either side can temporarily disturb the requirement established by the other side.
  4. GPI stabilizes when the policy is greedy with respect to its own evaluation function.
  5. At stabilization, the Bellman optimality equation holds, so the policy-value pair is optimal.

Key Takeaways

  • Generalized Policy Iteration is the repeated interaction of policy evaluation and policy improvement.
  • Improvement can make the current value function inaccurate by changing the policy it should describe.
  • Evaluation can remove a policy's greediness by changing the value function against which greediness is judged.
  • The processes stabilize when the policy is greedy with respect to its own evaluation function.
  • That stabilized pair satisfies the Bellman optimality equation and is optimal.