Computing v* with Policy Iteration
Policy iteration is applicable to action values, not only state values.
The Translation Problem
The central challenge is not to calculate a particular number. It is to translate the policy-iteration idea used for computing v* into an action-value form. The exercise asks what policy iteration would look like when q-values, rather than state values, are the main quantities being represented and updated.
The target of the action-value exercise is q*, in analogy with the state-value target v*.
Two Targets, Two Representations
The notation v* identifies the state-value target named in the original exercise. The notation q* identifies the corresponding action-value target. The important distinction is the object being represented: a state-value formulation is organized around values associated with states, while the action-value formulation is organized around values associated with state-action choices. Therefore, adapting the algorithm requires more than replacing one symbol with another. The representation, the iteration process, and the final result must all be expressed in action-value language.
| Question | State-value version | Action-value version |
|---|---|---|
| Named target | v* | q* |
| Main representation | Values associated with states | Values associated with state-action pairs |
| Design task | Use policy iteration to compute v* | Reformulate policy iteration to compute q* |
| Final check | The procedure produces the state-value target | The procedure produces the action-value target |
The action-value version changes the main representation and the stated result, not merely the symbol used in the final line.
The Policy-Iteration Cycle
A complete action-value formulation can be organized as a repeated cycle. First, choose an action-value representation for the current policy. Next, perform policy evaluation in that representation so the current policy has corresponding q-values. Then perform policy improvement using those q-values to determine the next policy. Continue the evaluation and improvement process until the policy-iteration procedure has reached its final result. The result is reported as q*.
The algorithm is complete only when it connects three things: what is represented, how policy iteration changes that representation, and how the final q* result is obtained.
A Complete Algorithm Blueprint
- Specify the action-value representation. State clearly that the quantities being stored or computed are q-values associated with state-action pairs.
- Choose or receive an initial policy. The policy supplies the current action choices whose action values will be evaluated.
- Evaluate the current policy using action values. The evaluation stage must produce q-values for the current policy rather than leaving the solution in state-value notation.
- Improve the policy using the evaluated q-values. The improvement stage uses the action-value representation to revise the policy's action choices.
- Repeat action-value policy evaluation and policy improvement as required by the policy-iteration procedure.
- When the iteration has reached its final policy result, identify the corresponding action-value result as q*.
- State the output explicitly. A complete answer should say that the algorithm computes q*, not v*.
A Worked Design Example
Turning an Incomplete Prompt into an Algorithm
An exercise asks how policy iteration can be defined for action values and requests a complete algorithm for computing q*, but it does not provide a numerical environment.
Identify the target: The requested result is q*, not v*. This determines that the solution must be written as an action-value algorithm.
Choose the representation: Declare that the algorithm represents values associated with state-action pairs. This prevents the explanation from silently reverting to a state-value formulation.
Describe evaluation: Say that policy evaluation produces q-values for the current policy. The current policy is therefore evaluated through its action-value representation.
Describe improvement: Say that the evaluated q-values are used in the policy-improvement stage to produce the next policy.
Close the loop: Repeat the action-value evaluation and improvement stages, then identify the final action-value result as q*.
Check numerical sufficiency: Do not invent numerical values. The supplied material contains no particular environment, action set, reward description, initial policy, or earlier v* algorithm.
The answer is a complete algorithmic blueprint for action-value policy iteration, but not a numerical q* table.
This example demonstrates the level of completeness expected from an algorithm-design answer. It does not pretend that an unspecified environment has numerical q-values. Instead, it makes the representation, iteration stages, target, and missing inputs explicit.
Missing Inputs for Calculation
The supplied material is enough to design the requested action-value version conceptually, but it is not enough to calculate numerical q* values. The source specifically notes that no particular environment, action set, reward description, initial policy, or earlier v* algorithm is supplied.
Common Specification Errors
Treating q* as if it were simply another spelling of v*
The source identifies the main task as reformulating the entire policy-iteration process using action values.
Fix:
Make the representation, evaluation stage, improvement stage, and final output action-value based.Giving only the target name
A complete algorithm must connect the representation, the iteration process, and the q* result.
Fix:
Include the action-value representation, policy evaluation, policy improvement, repetition, and final result.Inventing a numerical environment
The source provides none of the particular numerical information needed for a calculation.
Fix:
State the missing information and give a general algorithm blueprint instead.Leaving policy evaluation and improvement disconnected
The policy-iteration process must show how evaluation produces action values and how improvement uses them to update the policy.
Fix:
Present the stages as a cycle: represent, evaluate, improve, repeat, and report q*.
Practice: Audit an Algorithm
A proposed answer says: Start with a policy, evaluate its state values, improve the policy, and repeat until the result is v*. The exercise, however, asks for the action-value version and the target q*. Rewrite the answer as an action-value algorithm.
Hints
- Replace the state-value representation with an action-value representation.
- Name what policy evaluation produces in the adapted procedure.
- Explain what policy improvement uses to form the next policy.
- End by identifying the result as q*, while avoiding invented numerical details.
A strong revision should mention q-values throughout the process, not only in the final line.
Final Checklist
- The action-value target named by the exercise is q*, in analogy with v*.
- The central task is to reformulate policy iteration so that action values are the main quantities represented and updated.
- A complete algorithm connects the action-value representation, policy evaluation, policy improvement, repeated iteration, and the final q* result.
- The supplied material does not support a numerical q* calculation because it omits a particular environment, action set, reward description, initial policy, and earlier v* algorithm.
- When numerical inputs are missing, provide a precise algorithmic blueprint rather than inventing values.
Key Takeaways
- Policy iteration is applicable to action values as well as state values.
- The action-value exercise targets q*, while the state-value target is v*.
- The adaptation must use action-value language through representation, evaluation, improvement, and output.
- No numerical q* result can be computed without the environment and other required problem information.