Policy Iteration for Optimal Decision Making
Policy iteration is applicable to action values, not only state values.
From v* to q*
Policy iteration is not restricted to state values. The central task is to reformulate the policy-iteration process so that action values are the quantities represented, computed, and updated throughout. The target of this action-value version is q*, in analogy with the state-value target v* named in the source exercise.
The important change is not merely renaming v* as q*. A complete solution must specify what is stored, how policy evaluation and policy improvement operate on that representation, and how the final q* result is obtained.
The Iteration Loop
The action-value algorithm can be organized by adapting the policy-iteration approach used for v*. First choose or receive a policy and establish an action-value representation. Then evaluate that policy using action values. Next improve the policy using the evaluated action values. Repeat evaluation and improvement until the policy and its action-value representation reach the optimal result. The final action-value result is identified as q*.
- Represent the quantities to be computed as action values for state-action pairs.
- Start with a policy, whether supplied by the problem or selected as the initial policy for the algorithm.
- Evaluate the current policy using the action-value representation.
- Improve the policy by using the evaluated action values to select the preferred action in each state.
- Repeat evaluation and improvement while the policy is still changing.
- When the policy-iteration process has reached its stable result, report the corresponding action-value result as q*.
A Representation to Track
The representation determines what the algorithm can evaluate and improve. For v*, the central quantity is organized by state. For q*, the central quantity must distinguish actions as well as states. That means the action-value representation must make it possible to refer to the value associated with a particular state-action pair and to compare the available action values when improving the policy.
This organization is why an action-value solution cannot simply copy a state-value solution and change the final symbol. The algorithm must retain action-specific information throughout policy evaluation and policy improvement. Otherwise, there would be no action-value basis for revising the policy.
Abstract Walkthrough
Consider a generated abstract setting with two states and several available actions. No rewards, transitions, numerical values, initial policy, or discount information are supplied here. The purpose of this example is therefore to trace the structure of the algorithm, not to calculate q* numerically.
Tracing the Action-Value Version
Organize policy iteration for a setting in which each state has more than one available action, with q* as the desired result.
Represent: Create an action-value representation that distinguishes the value associated with each state-action pair. This is the quantity that the algorithm will evaluate and update.
Initialize: Begin with a policy. The policy identifies the current action choice associated with each state.
Evaluate: Evaluate the current policy using action values. The result describes the action values associated with following the current policy.
Improve: Use the evaluated action values to revise the policy. In each state, the action-value information provides the basis for selecting the preferred action.
Repeat: Return to policy evaluation after an improvement. Continue alternating the two stages while the policy is not yet stable.
Report: When the iteration reaches its stable optimal result, report the corresponding action-value function as q*.
The complete action-value policy-iteration design connects the q representation, repeated evaluation and improvement, and the final q* result.
The example is complete as an algorithmic outline but deliberately not numerical. The source material does not specify an environment or numerical inputs from which particular q* values could be calculated.
Inputs Before Calculation
Before attempting a numerical q* computation, check whether the problem supplies the information needed to define and run the iteration. The supplied source does not provide a particular environment, action set, reward description, initial policy, or earlier v* algorithm. Consequently, it supports an algorithm-design answer but not a numerical q* result.
Treating q* as though it were only a renamed v*.
The action-value formulation must use action values as the main quantities being represented and updated.
Fix:
Describe the representation, evaluation stage, improvement stage, and final result in action-value terms.Giving only the policy-improvement step.
A complete algorithm must connect representation, iteration, and the q* result.
Fix:
Include both policy evaluation and policy improvement, together with the stopping and reporting stages.Inventing numerical q* values when the problem supplies no numerical environment.
The supplied material supports an algorithm-design answer, not a numerical calculation.
Fix:
State which inputs are missing and give the complete abstract algorithm instead.Stopping after policy evaluation.
Policy iteration requires the evaluation and improvement stages to alternate toward the optimal result.
Fix:
Improve the policy after evaluation and repeat until the process reaches its stable result.
Practice Check
Write a six-step outline for an action-value policy-iteration algorithm. Your outline should begin with the q representation, include policy evaluation and policy improvement, explain when the process repeats, and identify the final result. Then list the information you would request before attempting a numerical computation.
Hints
- Use action-value language rather than replacing q* with v*.
- Separate evaluating the current policy from improving the policy.
- Check for the environment, action set, reward description, initial policy, and discount information.
What do you think happens?
If an exercise gives no particular environment, action set, reward description, initial policy, or earlier v* algorithm, can it determine numerical q* values?
Reveal answer
Answer: No. It can support an algorithm-design answer, but the missing numerical information prevents a numerical q* calculation.
The source explicitly describes the supplied material as defining the scope of the exercise rather than providing a numerical calculation.
What to Remember
- Policy iteration applies to action values as well as state values.
- The action-value target is q*, while v* is the state-value target named in the source exercise.
- A complete q* algorithm must specify the action-value representation, policy evaluation, policy improvement, repetition, and final reporting.
- Policy improvement uses the evaluated action values to select the preferred action in each state.
- Numerical q* computation requires supplied problem information; the source material alone provides an algorithm-design task rather than numerical inputs.
Key Takeaways
- The action-value version of policy iteration is intended to compute q*.
- q* is organized around state-action pairs, whereas v* is the state-value target named in the source exercise.
- The algorithm alternates action-value policy evaluation and policy improvement.
- A valid solution connects representation, iteration, and the final q* result.
- Without an environment, actions, rewards, initial policy, and other required information, only the algorithmic structure—not numerical q* values—can be given.