Concepts / Dynamic Programming and Policy Iteration

Dynamic Programming and Policy Iteration

GPI is a coordinated process, not a choice between policy improvement and value estimation.

  • Programming

One Process, Two Improvements

Generalized policy iteration, or GPI, is not a choice between estimating values and improving a policy. It is a coordinated process that keeps both an approximate policy and an approximate value function in play. The value function follows the current policy more closely, while the policy improves according to the current value function.

The central pattern is alternating improvement: estimate how well the current policy performs, then use that information to improve the policy.

The GPI Improvement Loop

estimate performanceimproveevaluate nextrepeatCurrent policyValue estimatehow well the policyperformsImproved policyuses the current valueestimateNext evaluationthe updated policy becomesthe target
What happens next as the value function is evaluated or improved, then used to improve the policy, and the updated policy becomes the next evaluation target?

The loop has two connected activities. First, the value function is improved so that it follows the current policy more closely. Second, the policy is improved using the current value function. After the policy changes, value estimation has a new policy to follow, so the process continues with a new evaluation target.

  1. Keep an approximate policy and an approximate value function available together.
  2. Estimate how well the current policy performs.
  3. Use the current value information to improve the policy.
  4. Treat the updated policy as the target for the next value-estimation phase.
  5. Repeat the connected improvements rather than expecting one update to produce an optimal policy.

A Policy and Its Value Estimate

guides estimationguides improvementupdatesupdatesApproximate policycurrent behaviorApproximate valuefunctionestimated performancePolicy improvementValue improvement
How are the approximate policy and approximate value function connected, and which one supplies information to improve the other?

The policy and value function have different roles, but they cannot be treated as unrelated objects. The current policy supplies the behavior whose performance is estimated. The resulting value information then supplies guidance for improving the policy. GPI maintains both approximations because each one provides information needed by the other.

Monte Carlo Information for Control

supports estimationinforms improvementgenerates performance evidenceSampled experienceobserved returnsValue estimateestimated policyperformanceApproximate policycandidate behaviorImproved policyapproximate optimal policy
How does sampled experience move from observed returns into value estimates and then into an improved policy?

Monte Carlo methods provide an estimation approach for control problems that seek approximate optimal policies. Their role in this pattern is not to produce an optimal policy in one immediate update. Instead, sampled experience contributes to an estimate of how well the current policy performs. That estimate can then be used to improve the policy.

Monte Carlo estimation contributes to control by supplying performance information for policy improvement. It is part of the repeating estimate-then-improve pattern.

Tracing One GPI Cycle

A generated abstract cycle

Trace what happens when GPI begins with an approximate policy and an approximate value function.

Starting point: The process has an approximate policy and an approximate value function. Neither is treated as permanently fixed.

Value improvement: Information about the current policy is used to make the value function follow that policy more closely.

Policy improvement: The updated value information is used to improve the current policy.

New target: Because the policy has changed, the value function must now represent a different policy. The next evaluation phase therefore has a changed target.

Continuation: The process repeats, with value-function improvement and policy improvement continuing as connected parts of one control process.

The cycle maintains and improves both approximations together rather than selecting only one of the two updates.

What do you think happens?

After the policy improves, does the existing value function still have exactly the same evaluation target?

  • Yes, because only the policy changed
  • No, because the value function must now represent a different policy
Reveal answer

Answer: No, because the value function must now represent a different policy.

The policy change creates a new target for value estimation. This is one reason GPI involves moving targets.

Why the Targets Move

is evaluated bysupports improvementcreates new targetPolicy Avalue function evaluates itPolicy Bimproved using valueinformationValue estimate Afollows Policy AValue estimate Bfollows Policy B
How can the value function and policy each change the target the other is trying to approximate while the overall process still moves toward optimality?

The two updates do not operate against completely fixed targets. When the value function changes, the policy is being judged using changed evaluation information. When the policy changes, the value function must now represent a different policy. Each update therefore changes what the other update is trying to approximate.

One Coordinated Control Process

evaluateinformupdatecontinueApproximate policyValue estimationfollows current policyPolicy improvementuses current value functionUpdatedapproximationscontinue the process
How do policy improvement and value estimation operate as coordinated parts of one control process instead of mutually exclusive alternatives?

It is misleading to ask whether GPI uses policy improvement or value estimation. It uses both in an organized relationship. Estimation provides information about the current policy, and policy improvement uses that information to produce a better approximation. The updated policy then becomes the subject of further estimation.

Part of GPIWhat it followsWhat it changes
Value-function improvementThe current policyThe value estimate
Policy improvementThe current value functionThe policy

Mistakes About Policy Iteration

  • Treating Monte Carlo control as one update that immediately produces an optimal policy.

    The source describes a repeating pattern: estimate how well the current policy performs, then use that information to improve the policy.

    Fix: Think in terms of repeated connected estimation and policy-improvement steps.

  • Choosing between value estimation and policy improvement as if only one belongs in GPI.

    GPI is defined as a coordinated process that maintains an approximate policy and an approximate value function together.

    Fix: Track the two updates as complementary parts of one control process.

  • Assuming that the value function has a permanently fixed target.

    When the policy changes, the value function must represent a different policy.

    Fix: Recognize that each policy update creates a new target for value estimation.

  • Interpreting moving targets as proof that GPI cannot approach optimality.

    Although the updates work against each other to some extent, the source states that they still approach optimality together.

    Fix: Understand moving targets as a feature of the coordinated process, not as evidence that the process stops.

Check Your Understanding

MEDIUM

Explain the GPI cycle in your own words. Your explanation should identify the two approximations being maintained, state what the current policy contributes to value estimation, state what the current value function contributes to policy improvement, and explain why a policy change creates a new value-estimation target.

Hints
  • Begin with the current approximate policy.
  • Describe how its performance is estimated.
  • Connect that estimate to policy improvement.
  • Finish by explaining why the updated policy changes the next evaluation target.
MEDIUM

A learner says, “Because the policy and value function keep changing, GPI is just switching randomly between two methods.” Correct the statement using the ideas of coordination, moving targets, and approaching optimality together.

Hints
  • GPI is a coordinated process, not a choice between methods.
  • Each part uses information from the other.
  • Moving targets do not prevent the two parts from approaching optimality together.

The Essential Pattern

  1. GPI jointly maintains an approximate policy and an approximate value function.
  2. Monte Carlo estimation contributes performance information that can help improve an approximate policy.
  3. Value-function improvement follows the current policy more closely, while policy improvement uses the current value function.
  4. The updates create moving targets because changing either approximation changes what the other is trying to represent.
  5. Despite those moving targets, the coordinated process approaches optimality together.

Key Takeaways

  • Generalized policy iteration is a coordinated process rather than a choice between value estimation and policy improvement.
  • It keeps an approximate policy and an approximate value function active at the same time.
  • Monte Carlo estimation helps control by estimating how well the current policy performs, after which the policy can be improved.
  • The policy and value function create moving targets for each other, but their connected updates still approach optimality together.