Concepts / Optimal Policies and Value Functions

Optimal Policies and Value Functions

Evaluation and improvement impose different requirements on a policy-value pair.

  • Programming

A Moving Target

Policy evaluation and policy improvement both work on a policy-value pair, but they ask for different things. Evaluation asks the value function to describe the current policy accurately. Improvement asks the policy to become greedy with respect to the current value function. Because each process changes one side of the pair while relying on the other, progress by one process can temporarily disturb the condition the other process was establishing.

requires consistencyrequires greedinessPolicy evaluationValue matches policyPolicy-value pairShared objectPolicy improvementPolicy is greedy
How can policy evaluation push a policy-value pair toward consistency while policy improvement pushes it toward greater action value?

The First Disturbance

A Policy Change Outruns Its Values

Consider a policy-value pair in which the value function has been made consistent with the current policy. Policy improvement then examines that value function and changes the policy to make it greedy with respect to those values.

1. Start with a fitted pair: The value function describes the current policy, so the evaluation requirement is being met for that pair.

2. Improve the policy: The policy changes in response to the current value function. This is the improvement process acting on the pair.

3. Reconsider the old values: The value function has not yet been changed to account for the new policy. It therefore no longer necessarily describes the changed policy accurately.

4. Evaluate again: Evaluation must respond to the policy change by bringing the value function into consistency with the changed policy.

The policy change can make the current value function inaccurate for that policy, so improvement can create work for evaluation.

value describesneeds evaluationPolicy ACurrent policyPolicy BChanged policyValue VFits Policy AValue VNot yet refitted
What changes in the value estimates immediately after the policy changes, and why do those estimates no longer describe the new policy?

What do you think happens?

After policy improvement changes the policy but before another evaluation step, what is the safest description of the existing value function?

  • It is guaranteed to describe the new policy accurately
  • It may no longer describe the new policy accurately
  • It disappears and must be replaced from nothing
Reveal answer

Answer: It may no longer describe the new policy accurately

Improvement changes the policy using the current value function. Until evaluation accounts for that changed policy, the old value function can be inaccurate for the new policy.

Why the Requirements Compete

The processes compete in their immediate effects because each one can undo the other's current achievement. If evaluation makes the value function consistent with a policy, the policy may no longer be greedy with respect to the newly consistent values. Conversely, if improvement makes the policy greedy with respect to the current values, the policy may change enough that the current value function no longer fits it. The competition is therefore local and temporary: it describes what can happen immediately after one process acts, not the final purpose of the method.

value must match policygreediness can changePolicy AGreedy for current valuesEvaluationRefits valuesPolicy AMay not be greedy
How can updating values to accurately reflect a policy cause that policy to stop being greedy with respect to those same values?

The Alternating Search

Generalized Policy Iteration can be viewed as an alternating cycle. Improvement reads the current value function and changes the policy toward greediness. That change can disturb the value function's fit to the policy. Evaluation then works to make the value function match the changed policy. The updated values can in turn affect whether the policy remains greedy, so improvement may have another reason to act. The processes repeatedly influence one another while searching for a shared solution.

read valueschanges policyrequires refittingupdates valuesmay trigger another stepPolicy-value pairCurrent stateImprovementPolicy becomes greedyChanged policyOld values may not fitEvaluationValues match policyUpdated pairContinue the search
What happens next as evaluation and improvement alternate, and how does each step affect the other?

Following One Complete Turn

Track an abstract policy-value pair through improvement and evaluation without assigning names or numerical values to particular actions.

Initial condition: The value function has been adjusted to match the current policy.

Improvement acts: The policy is changed toward being greedy with respect to the current value function.

Immediate consequence: Because the policy changed, the existing value function may no longer accurately describe that policy.

Evaluation acts: The value function is adjusted so that it matches the changed policy more closely.

Next interaction: After the values change, the policy's greediness must again be considered. The two processes remain linked.

Each process responds to the current condition created partly by the other process.

The Shared Destination

The long-run goal is not to keep evaluation and improvement permanently opposed. They cooperate toward one joint solution: an optimal value function together with an optimal policy. At that destination, the value function satisfies the consistency requirement associated with the policy, and the policy satisfies the improvement requirement associated with the value function. The two processes may disturb one another locally, but their interaction is what guides the search toward this shared result.

meetsmeetsValue consistencyEvaluation requirementOptimal pairOptimal policy and valuefunctionPolicy greedinessImprovement requirement
Where do the consistency requirement from evaluation and the greediness requirement from improvement meet in the optimal policy-value pair?

Two Constraints, One Intersection

A useful geometric analogy treats evaluation and improvement as two separate constraints in a two-dimensional space. One constraint represents the demand that the value function match the policy. The other represents the demand that the policy be greedy with respect to the value function. A policy-value pair that satisfies only one constraint is incomplete. Their intersection represents the shared solution sought by the two processes. This picture is conceptual rather than literal: the actual geometry of the problem is much more complicated than two lines.

shares a solution withshares a solution withEvaluationconstraintValue matches policyIntersectionOptimal policy-value pairImprovementconstraintPolicy is greedy
How do two separate constraints restrict the possible policy-value pairs, and what does their intersection represent?

Common Reasoning Errors

  • Assuming that a greedy policy change leaves the value function accurate automatically.

    A policy change can invalidate the value function's fit to the policy.

    Fix: Treat the changed policy as a reason for evaluation to act again.

  • Assuming that evaluation permanently guarantees greediness.

    Consistency with a policy and greediness with respect to a value function are different requirements.

    Fix: Check both requirements separately.

  • Treating temporary disruption as the final outcome.

    The processes compete in their immediate effects but cooperate toward an optimal value function and optimal policy.

    Fix: Interpret the disruption as part of an alternating search for a joint solution.

  • Reading the two-constraint analogy as a literal geometric description.

    The source uses that picture conceptually; the actual geometry is much more complicated.

    Fix: Use the analogy to organize the requirements, not to claim that the real problem has exactly that geometry.

Check Your Model

MEDIUM

A policy-value pair has been evaluated, so the value function matches the current policy. Improvement then changes the policy to make it greedy with respect to the current values. Explain why evaluation may need to act again, and state what joint destination the two processes are seeking.

Hints
  • Ask whether the value function was computed for the old policy or the changed policy.
  • Keep consistency and greediness as two separate requirements.
  • Name the shared destination rather than describing only the temporary disruption.

A Strong Answer

Explain the effect of improvement and the long-run purpose of the alternating processes.

Identify the change: Improvement changes the policy while using the current value function.

Identify the mismatch: The current value function may no longer accurately describe the changed policy.

Identify the response: Evaluation must work to make the value function consistent with the changed policy.

Identify the destination: The processes seek an optimal value function together with an optimal policy.

Improvement can create a temporary evaluation mismatch, but both processes cooperate in the long run toward one optimal policy-value pair.

Key Takeaways

  1. Evaluation asks the value function to match the current policy.
  2. Improvement asks the policy to become greedy with respect to the current value function.
  3. Changing the policy can make the existing value function inaccurate for that changed policy.
  4. Making values consistent with a policy can change whether that policy is greedy.
  5. The two processes compete locally but cooperate toward an optimal value function and an optimal policy.

Key Takeaways

  • Policy evaluation and policy improvement impose different requirements on a policy-value pair.
  • Improvement can disturb value accuracy, while evaluation can disturb policy greediness.
  • Generalized Policy Iteration alternates these processes as they respond to one another.
  • The two-constraint analogy represents consistency and greediness meeting at an optimal policy-value pair.
  • Temporary competition between the processes supports long-run cooperation toward an optimal solution.