Concepts / Policy Evaluation and Policy Improvement

Policy Evaluation and Policy Improvement

Optimality describes the highest-value policy, not the computational effort required to find it.

  • Programming

The Gap Between Best and Feasible

An optimal policy is the policy that achieves the highest value function. This definition describes the strongest possible decision rule, but it does not promise that an agent can calculate that policy in the time available. Even a complete and accurate description of the environment may leave optimal policy computation too demanding.

constrainsOptimal policyhighest valueFeasible policycomputable in timeAvailable computationtime and processing
What is the difference between the policy with the highest possible value and a policy an agent can compute with available resources?

Computation Inside a Time Step

Knowing how the environment behaves is only one part of decision-making. The agent must also use that information to evaluate possibilities and choose a policy quickly enough. A time step creates a practical boundary: computation that is unfinished when a decision is required cannot improve that immediate decision.

Chess illustrates the distinction. Large, specialized computers can play strongly, but substantial computational resources still do not guarantee that they can compute an optimal move. The lesson is not that useful chess policies are impossible. The lesson is that a carefully defined environment and powerful computation do not guarantee an optimal policy.

beginsmust finish beforerequiresTime step beginscomputation availableEvaluatepossibilitiesuse model informationDecision requiredtime step boundaryChoose actionavailable result
How does limited computation before an action is chosen restrict the planning an agent can use?

Evaluating a Fixed Policy

Policy evaluation measures a fixed policy by computing the value functions associated with following it. It does not change the policy. Instead, it determines what that current decision rule achieves.

Evaluation is usually iterative. Values are repeatedly updated until the updates stop changing them. Repeating the process matters because the value of a state depends on what can happen later when the policy continues to be followed. Evaluation therefore turns a policy into value information that can be used in the next stage.

supplies valuessupplies outcomeupdatescontinues sweeprepeatsSuccessor statescurrent value estimatesRewardoutcome informationFull backupmodel probabilitiesState valueupdated estimateRepeated sweepuntil values stop changing
How do rewards and successor-state values flow backward through repeated backups to estimate the value of following a policy?

What evaluation changes

Suppose a policy already specifies what to do in every state. What does policy evaluation do with that policy?

Keep the policy fixed: Evaluation does not replace the current decision rule with another one.

Update values: For each state, use rewards and the values of possible successor states to update the estimate of following the fixed policy.

Repeat the updates: Continue the backups through repeated sweeps until the values stop changing.

The result is a value function describing what the fixed policy achieves.

Full Backups and Bellman Relationships

A full backup is the local operation behind a sweep. Choose one state, then use the values of all possible successor states together with their probabilities of occurring to update the value of the chosen state. Full means that the update considers every successor represented by the model rather than only one observed successor.

A Bellman equation describes how a state's value relates to the values of its successor states. A backup turns that relationship into an update operation. Applying backups across the state set creates a sweep. If repeated backups eventually produce no further changes, the resulting values satisfy the relevant Bellman equation.

Improving the Decision Rule

Policy improvement uses the value function of the current policy to construct a better policy. Evaluation measures the current rule; improvement uses that measurement to alter the rule.

At a state, improvement uses the available value information to compare the consequences of possible decisions and select a policy that is better informed by those values. The policy is changed only in this improvement stage, not during evaluation.

presentscompared usinginformsconstructsCurrent statedecision pointPossible actionscandidate choicesValue informationsuccessor outcomesBetter choiceimproved decisionImproved policyupdated rule
How does the value function compare possible actions at a state and select the action that produces a better policy?

What improvement changes

A current policy has been evaluated, so its value function is available. What happens next?

Use the evaluation: Treat the current policy's value function as information about what the existing decision rule achieves.

Consider alternatives: Use the known environment dynamics and value information to compare possible decisions at a state.

Construct a candidate: Select the decision supported by the comparison and incorporate it into an improved policy.

The policy changes because improvement uses the current value function to produce a better-informed decision rule.

The Policy-Iteration Loop

Policy iteration alternates two jobs. Begin with a policy, evaluate it, improve it, and then evaluate the resulting policy again. The loop continues until improvement no longer changes the policy. For a finite Markov decision process with complete knowledge of the process, the source describes policy iteration as a method that can reliably compute an optimal policy and its value functions.

startvalues availablecheck resultpolicy changedno changeCurrent policystarting rulePolicy evaluationcompute value functionsPolicy improvementconstruct better rulePolicy changedcontinue loopStable policystop improving
How does the agent alternate between evaluating the current policy and improving it until the policy stops changing?

What do you think happens?

After policy improvement changes the current policy, should the agent stop or evaluate the new policy again?

  • Stop immediately
  • Evaluate the new policy again
  • Discard the new policy
Reveal answer

Answer: Evaluate the new policy again.

Policy iteration repeatedly evaluates the current policy and then improves it. A changed policy becomes the next policy to evaluate.

Policy Iteration and Value Iteration

Policy iteration and value iteration are distinct dynamic programming methods. Both are used to compute optimal policies and value functions for finite Markov decision processes when the model is completely known. Policy iteration is defined by its explicit alternation between evaluating a policy and improving it. Value iteration follows a different route rather than using that same two-stage loop.

MethodWhat is explicitRole of value information
Policy iterationA policy-evaluation stage followed by a policy-improvement stageEvaluate the current policy, then use its values to construct a better policy
Value iterationA distinct dynamic programming route rather than the same two-stage loopAims to compute optimal policies and value functions

The central distinction is whether the method explicitly alternates evaluation and improvement as separate stages.

evaluatethen improverepeataims towardCurrent policyexplicit ruleEvaluatepolicy valuesImprovebetter ruleValue iterationdistinct routeOptimal policy andvaluestarget
What changes at each update in policy iteration compared with value iteration?

Common Reasoning Errors

  • Treating an accurate environment model as a guarantee of an optimal policy

    The computation required to evaluate possibilities and compare policies may be too large to complete before the decision is required.

    Fix: Separate information about the environment from the computation needed to use that information.

  • Thinking policy evaluation changes the policy

    Evaluation computes the value functions for a fixed policy; improvement is the stage that constructs a better policy.

    Fix: Keep the policy fixed during evaluation, then use its values during improvement.

  • Confusing one backup with a completed evaluation

    Policy evaluation is typically iterative, with repeated updates until the values stop changing.

    Fix: Think in terms of repeated backups and sweeps through the state set.

  • Calling value iteration the same as policy iteration

    Value iteration is presented as a separate method rather than the explicit two-stage loop of policy iteration.

    Fix: Identify the explicit evaluation-then-improvement cycle before calling a method policy iteration.

Check Your Understanding

MEDIUM

Explain the following sequence in your own words: start with a policy, evaluate it, improve it, evaluate the new policy, and stop when improvement no longer changes the policy. Then explain why the same sequence cannot be described as value iteration merely because both methods target optimal policies and value functions.

Hints
  • State what remains fixed during evaluation.
  • State what the value function contributes during improvement.
  • Focus on the explicit stages that define policy iteration.
EASY

A decision must be made before a large computation finishes. What does this reveal about the difference between an optimal policy and a feasible policy?

Hints
  • Optimality describes value.
  • Feasibility depends on completing computation in time.
  • Use the time step as the practical boundary.

Key Takeaways

  1. An optimal policy is the policy with the highest value function, but optimality does not imply that an agent can compute the policy in practice. Complete environment knowledge is not enough when the required computation exceeds what can be completed within a time step. Policy evaluation computes the value functions of a fixed policy through repeated backups. Policy improvement uses those values to construct a better policy. Policy iteration alternates evaluation and improvement until the policy stops changing, while value iteration is a separate dynamic programming method. Full backups use all model-represented successor states and their probabilities, turning Bellman relationships into repeated value updates.

Key Takeaways

  • Optimality identifies the highest-value policy, not the amount of computation needed to find it.
  • An accurate model can still be difficult to use when the available computation cannot finish before a decision is required.
  • Policy evaluation measures a fixed policy by repeatedly updating its value functions.
  • Policy improvement uses those values to construct a better policy.
  • Policy iteration alternates evaluation and improvement, whereas value iteration follows a distinct route.