Concepts / Policy Iteration

Policy Iteration

GPI describes a feedback relationship, not one fixed algorithm.

  • Programming

The Two-Way Improvement Loop

Generalized policy iteration, or GPI, is a feedback relationship between two processes. Policy evaluation updates the value function so that it better represents the current policy. Policy improvement updates the policy so that it becomes greedy with respect to the current value function. GPI is therefore a broad idea about how these processes interact, not one fixed algorithm or schedule.

evaluateproducesguidechangesCurrent policyPolicy evaluationupdates value functionCurrent valuefunctionPolicy improvementupdates policy
How do policy evaluation and policy improvement feed into one another over repeated iterations?

The dependency runs in both directions: improvement responds to the value function, while evaluation responds to the policy. Neither process is the complete explanation of GPI by itself.

What Changes in Each Process

During policy evaluation, the policy is the current policy being assessed, while the value function is updated to become consistent with that policy. Evaluation does not choose a new policy. During policy improvement, the current value function is used to change the policy so that the policy becomes greedy with respect to that value function. The value function is the reference for the improvement step.

evaluatesguidesPolicycurrent policyPolicyupdatedValue functionupdatedValue functioncurrent value function
Which part stays fixed and which part changes during policy evaluation, and how is that reversed during policy improvement?

Tracing One GPI Round

Follow what happens when a current policy is evaluated and then improved.

Start: Begin with a policy and a value function. The policy is the behavior currently being assessed.

Evaluate: Keep the current policy as the policy under assessment, and update the value function so that it better reflects that policy.

Improve: Use the updated value function to change the policy toward choices that are greedy with respect to that value function.

Continue: Because the policy may have changed, the value function may no longer be the correct evaluation of the new policy. Evaluation and improvement therefore continue to affect one another.

One GPI round changes the value function during evaluation and changes the policy during improvement. The round is not complete convergence unless both processes eventually stop producing changes together.

Scheduling the Interaction

GPI does not prescribe one fixed schedule. Policy iteration alternates the processes at a coarse level: one process completes before the other begins. Value iteration uses only a single iteration of policy evaluation between policy improvements. Asynchronous dynamic programming can interleave the processes even more finely, including updating a single state in one process before returning to the other.

after completionafter one iterationafter a fine-grained updateEvaluationcomplete processEvaluationone iterationState updatefine interleavingImprovementthen complete processImprovementthen repeatOther processreturns quickly
When do evaluation and improvement occur, and are they applied after every step, after an episode, or across many states at once?
MethodEvaluation and improvement scheduleGranularity
Policy iterationEvaluation and improvement alternate, with one process completing before the other beginsCoarser alternation
Value iterationA single iteration of policy evaluation occurs between policy improvementsFiner interleaving
Asynchronous dynamic programmingThe processes can interleave at an even finer level, including a single-state updateVery fine interleaving

The methods differ in timing and granularity, not in the basic roles of the two processes. Evaluation still makes the value function consistent with the policy, and improvement still makes the policy greedy with respect to the value function.

Why the Processes Pull Apart

Evaluation and improvement can temporarily disturb one another. Suppose evaluation has made the value function consistent with the current policy. Improvement may then change the policy because the current value function indicates better choices. Once the policy changes, the old value function may no longer be the correct evaluation of that changed policy. Evaluation must respond again.

evaluate current policyimprove from current valuepolicy change may require reevaluationwhen no further change remainsPolicy and valuenot jointly alignedValue functionmade consistentPolicymade greedyPolicy-value pairneither process changes it
What changes from one iteration to the next, and how do the policy and value function eventually reach a stable pair?

This apparent conflict is not a failure of GPI. It is the feedback mechanism doing its work. Improvement exposes whether the policy can be made better according to the current value function. Evaluation exposes whether the value function still correctly represents the policy after that policy changes.

What do you think happens?

After policy improvement changes the policy, is the old value function automatically guaranteed to remain the correct evaluation of the new policy?

  • Yes, because improvement only changes the policy in a helpful direction
  • No, because changing the policy can make the previous value function incorrect for that new policy
Reveal answer

Answer: No, because changing the policy can make the previous value function incorrect for that new policy.

The value function was made consistent with the earlier policy. When improvement changes the policy, evaluation may be needed again to make the value function consistent with the changed policy.

The Joint Stopping Condition

GPI stabilizes when policy evaluation and policy improvement both stop producing changes. This means the value function is consistent with the policy, and the policy is greedy with respect to that value function. The decisive condition is therefore not that one process finishes while the other remains unfinished. Both requirements must hold together.

evaluation no longer changes valueimprovement no longer changes policyimpliesconfirmsValue functionconsistent with policyJoint stabilizationneither process changesBellman optimalityequationcondition holdsOptimal pairpolicy and value functionPolicygreedy with respect tovalue
How does a policy and value function that no longer change satisfy the Bellman optimality condition?

When the policy is greedy with respect to its own evaluation function, the Bellman optimality equation holds. The stabilized policy and value function are then optimal. This is stronger than saying that the policy was improved once or that the value function was evaluated once: the policy must be appropriate for the value function, and the value function must be the evaluation of that policy.

Mistakes in Reading GPI

  • Treating GPI as one fixed algorithm.

    GPI describes a feedback relationship, while different methods schedule the two processes at different levels of granularity.

    Fix: Separate the roles of evaluation and improvement from the schedule used to interleave them.

  • Saying that policy evaluation changes the policy.

    Evaluation updates the value function for the current policy.

    Fix: Associate evaluation with making the value function consistent with the current policy.

  • Saying that policy improvement evaluates the current policy.

    Improvement uses the current value function to update the policy.

    Fix: Associate improvement with making the policy greedy with respect to the current value function.

  • Assuming that one completed process proves optimality.

    The policy must also be greedy with respect to its own evaluation function.

    Fix: Check whether both processes stop producing changes together.

  • Treating temporary disagreement as evidence that GPI cannot converge.

    Changing one object can temporarily disturb the condition established by the other, and this interaction is part of GPI.

    Fix: Distinguish intermediate tension from the eventual joint solution.

When tracing an iteration, write down which object is being held conceptually current and which object is being updated. Then ask whether the update makes the value function more consistent with the policy or makes the policy more greedy with respect to the value function.

Check Your Understanding

MEDIUM

A learner says: “GPI has converged because policy evaluation stopped changing the value function, even though policy improvement would still change the policy.” Is this conclusion correct? Explain which stabilization condition is missing.

Hints
  • Identify what policy evaluation stopping means about the relationship between the value function and the policy.
  • Identify what policy improvement stopping means about the relationship between the policy and the value function.
  • Use the phrase joint stabilization in your explanation.

Answering the Stabilization Question

Determine whether GPI has converged when evaluation no longer changes the value function but improvement would still change the policy.

Interpret evaluation: The value function is currently consistent with the policy being evaluated.

Interpret improvement: Because improvement would still change the policy, the current policy is not yet greedy with respect to the current value function.

Apply the stopping condition: Both processes must stop producing changes. One process stopping is not enough.

GPI has not converged. The policy and value function have not reached their joint solution because the policy is still improvable with respect to the value function.

Key Takeaways

  1. Generalized policy iteration is a feedback relationship between policy evaluation and policy improvement, not one fixed algorithm.
  2. Evaluation updates the value function for the current policy, while improvement updates the policy using the current value function.
  3. Policy iteration, value iteration, and asynchronous dynamic programming differ in how finely they schedule the same two processes.
  4. Changing one object can temporarily disturb the condition established by the other, so convergence is a joint stabilization.
  5. When the policy is greedy with respect to its own evaluation function, the Bellman optimality equation holds and the policy-value pair is optimal.

Key Takeaways

  • GPI repeatedly connects policy evaluation and policy improvement.
  • Policy evaluation changes the value function for a current policy; policy improvement changes the policy using the current value function.
  • Different reinforcement learning methods vary the timing and granularity of this interaction.
  • The temporary conflict between the processes is expected because changing one object can make the other object outdated.
  • Joint stabilization occurs when the value function is consistent with the policy and the policy is greedy with respect to that value function; then the Bellman optimality equation holds and the pair is optimal.