Dynamic Programming and Policy Iteration
GPI is a coordinated process, not a choice between policy improvement and value estimation.
One Process, Two Improvements
Generalized policy iteration, or GPI, is not a choice between estimating values and improving a policy. It is a coordinated process that keeps both an approximate policy and an approximate value function in play. The value function follows the current policy more closely, while the policy improves according to the current value function.
The central pattern is alternating improvement: estimate how well the current policy performs, then use that information to improve the policy.
The GPI Improvement Loop
The loop has two connected activities. First, the value function is improved so that it follows the current policy more closely. Second, the policy is improved using the current value function. After the policy changes, value estimation has a new policy to follow, so the process continues with a new evaluation target.
- Keep an approximate policy and an approximate value function available together.
- Estimate how well the current policy performs.
- Use the current value information to improve the policy.
- Treat the updated policy as the target for the next value-estimation phase.
- Repeat the connected improvements rather than expecting one update to produce an optimal policy.
A Policy and Its Value Estimate
The policy and value function have different roles, but they cannot be treated as unrelated objects. The current policy supplies the behavior whose performance is estimated. The resulting value information then supplies guidance for improving the policy. GPI maintains both approximations because each one provides information needed by the other.
Monte Carlo Information for Control
Monte Carlo methods provide an estimation approach for control problems that seek approximate optimal policies. Their role in this pattern is not to produce an optimal policy in one immediate update. Instead, sampled experience contributes to an estimate of how well the current policy performs. That estimate can then be used to improve the policy.
Monte Carlo estimation contributes to control by supplying performance information for policy improvement. It is part of the repeating estimate-then-improve pattern.
Tracing One GPI Cycle
A generated abstract cycle
Trace what happens when GPI begins with an approximate policy and an approximate value function.
Starting point: The process has an approximate policy and an approximate value function. Neither is treated as permanently fixed.
Value improvement: Information about the current policy is used to make the value function follow that policy more closely.
Policy improvement: The updated value information is used to improve the current policy.
New target: Because the policy has changed, the value function must now represent a different policy. The next evaluation phase therefore has a changed target.
Continuation: The process repeats, with value-function improvement and policy improvement continuing as connected parts of one control process.
The cycle maintains and improves both approximations together rather than selecting only one of the two updates.
What do you think happens?
After the policy improves, does the existing value function still have exactly the same evaluation target?
Reveal answer
Answer: No, because the value function must now represent a different policy.
The policy change creates a new target for value estimation. This is one reason GPI involves moving targets.
Why the Targets Move
The two updates do not operate against completely fixed targets. When the value function changes, the policy is being judged using changed evaluation information. When the policy changes, the value function must now represent a different policy. Each update therefore changes what the other update is trying to approximate.
One Coordinated Control Process
It is misleading to ask whether GPI uses policy improvement or value estimation. It uses both in an organized relationship. Estimation provides information about the current policy, and policy improvement uses that information to produce a better approximation. The updated policy then becomes the subject of further estimation.
| Part of GPI | What it follows | What it changes |
|---|---|---|
| Value-function improvement | The current policy | The value estimate |
| Policy improvement | The current value function | The policy |
Mistakes About Policy Iteration
Treating Monte Carlo control as one update that immediately produces an optimal policy.
The source describes a repeating pattern: estimate how well the current policy performs, then use that information to improve the policy.
Fix:
Think in terms of repeated connected estimation and policy-improvement steps.Choosing between value estimation and policy improvement as if only one belongs in GPI.
GPI is defined as a coordinated process that maintains an approximate policy and an approximate value function together.
Fix:
Track the two updates as complementary parts of one control process.Assuming that the value function has a permanently fixed target.
When the policy changes, the value function must represent a different policy.
Fix:
Recognize that each policy update creates a new target for value estimation.Interpreting moving targets as proof that GPI cannot approach optimality.
Although the updates work against each other to some extent, the source states that they still approach optimality together.
Fix:
Understand moving targets as a feature of the coordinated process, not as evidence that the process stops.
Check Your Understanding
Explain the GPI cycle in your own words. Your explanation should identify the two approximations being maintained, state what the current policy contributes to value estimation, state what the current value function contributes to policy improvement, and explain why a policy change creates a new value-estimation target.
Hints
- Begin with the current approximate policy.
- Describe how its performance is estimated.
- Connect that estimate to policy improvement.
- Finish by explaining why the updated policy changes the next evaluation target.
A learner says, “Because the policy and value function keep changing, GPI is just switching randomly between two methods.” Correct the statement using the ideas of coordination, moving targets, and approaching optimality together.
Hints
- GPI is a coordinated process, not a choice between methods.
- Each part uses information from the other.
- Moving targets do not prevent the two parts from approaching optimality together.
The Essential Pattern
- GPI jointly maintains an approximate policy and an approximate value function.
- Monte Carlo estimation contributes performance information that can help improve an approximate policy.
- Value-function improvement follows the current policy more closely, while policy improvement uses the current value function.
- The updates create moving targets because changing either approximation changes what the other is trying to represent.
- Despite those moving targets, the coordinated process approaches optimality together.
Key Takeaways
- Generalized policy iteration is a coordinated process rather than a choice between value estimation and policy improvement.
- It keeps an approximate policy and an approximate value function active at the same time.
- Monte Carlo estimation helps control by estimating how well the current policy performs, after which the policy can be improved.
- The policy and value function create moving targets for each other, but their connected updates still approach optimality together.