Concepts / Computing v* with Policy Iteration

Computing v* with Policy Iteration

Policy iteration is applicable to action values, not only state values.

  • Programming

The Translation Problem

The central challenge is not to calculate a particular number. It is to translate the policy-iteration idea used for computing v* into an action-value form. The exercise asks what policy iteration would look like when q-values, rather than state values, are the main quantities being represented and updated.

The target of the action-value exercise is q*, in analogy with the state-value target v*.

represented byrepresented byv*state valuestatestate representationq*action valuestate-action pairaction-value representation
What is the difference between a value attached to a state and a value attached to a state-action pair?

Two Targets, Two Representations

The notation v* identifies the state-value target named in the original exercise. The notation q* identifies the corresponding action-value target. The important distinction is the object being represented: a state-value formulation is organized around values associated with states, while the action-value formulation is organized around values associated with state-action choices. Therefore, adapting the algorithm requires more than replacing one symbol with another. The representation, the iteration process, and the final result must all be expressed in action-value language.

QuestionState-value versionAction-value version
Named targetv*q*
Main representationValues associated with statesValues associated with state-action pairs
Design taskUse policy iteration to compute v*Reformulate policy iteration to compute q*
Final checkThe procedure produces the state-value targetThe procedure produces the action-value target

The action-value version changes the main representation and the stated result, not merely the symbol used in the final line.

The Policy-Iteration Cycle

A complete action-value formulation can be organized as a repeated cycle. First, choose an action-value representation for the current policy. Next, perform policy evaluation in that representation so the current policy has corresponding q-values. Then perform policy improvement using those q-values to determine the next policy. Continue the evaluation and improvement process until the policy-iteration procedure has reached its final result. The result is reported as q*.

current policyq-valuesupdated policyiteratefinishAction-valuerepresentationrepresent the currentpolicyPolicy evaluationproduce current-policyq-valuesPolicy improvementuse q-values to update thepolicyPolicy resultcontinue or finishq*action-value result
What happens in each policy-evaluation and policy-improvement step when policy iteration computes q* instead of v*?

The algorithm is complete only when it connects three things: what is represented, how policy iteration changes that representation, and how the final q* result is obtained.

A Complete Algorithm Blueprint

  1. Specify the action-value representation. State clearly that the quantities being stored or computed are q-values associated with state-action pairs.
  2. Choose or receive an initial policy. The policy supplies the current action choices whose action values will be evaluated.
  3. Evaluate the current policy using action values. The evaluation stage must produce q-values for the current policy rather than leaving the solution in state-value notation.
  4. Improve the policy using the evaluated q-values. The improvement stage uses the action-value representation to revise the policy's action choices.
  5. Repeat action-value policy evaluation and policy improvement as required by the policy-iteration procedure.
  6. When the iteration has reached its final policy result, identify the corresponding action-value result as q*.
  7. State the output explicitly. A complete answer should say that the algorithm computes q*, not v*.
policy defined overpolicy selectsstate-value viewaction-value viewstatesstate representationpolicyaction choicesvstate valuesactionsavailable choicesqaction values
How do a policy's action choices connect state values v with action values q?

A Worked Design Example

Turning an Incomplete Prompt into an Algorithm

An exercise asks how policy iteration can be defined for action values and requests a complete algorithm for computing q*, but it does not provide a numerical environment.

Identify the target: The requested result is q*, not v*. This determines that the solution must be written as an action-value algorithm.

Choose the representation: Declare that the algorithm represents values associated with state-action pairs. This prevents the explanation from silently reverting to a state-value formulation.

Describe evaluation: Say that policy evaluation produces q-values for the current policy. The current policy is therefore evaluated through its action-value representation.

Describe improvement: Say that the evaluated q-values are used in the policy-improvement stage to produce the next policy.

Close the loop: Repeat the action-value evaluation and improvement stages, then identify the final action-value result as q*.

Check numerical sufficiency: Do not invent numerical values. The supplied material contains no particular environment, action set, reward description, initial policy, or earlier v* algorithm.

The answer is a complete algorithmic blueprint for action-value policy iteration, but not a numerical q* table.

This example demonstrates the level of completeness expected from an algorithm-design answer. It does not pretend that an unspecified environment has numerical q-values. Instead, it makes the representation, iteration stages, target, and missing inputs explicit.

Missing Inputs for Calculation

The supplied material is enough to design the requested action-value version conceptually, but it is not enough to calculate numerical q* values. The source specifically notes that no particular environment, action set, reward description, initial policy, or earlier v* algorithm is supplied.

suppliessuppliessuppliesstartscomputesenvironmentparticular settingaction-value policyiterationevaluation and improvementq*numerical resultaction setavailable actionsreward descriptionreward informationinitial policystarting action choices
What model, rewards, discount factor, states, actions, and initial policy must be supplied before the algorithm can compute numerical q-values?

Common Specification Errors

  • Treating q* as if it were simply another spelling of v*

    The source identifies the main task as reformulating the entire policy-iteration process using action values.

    Fix: Make the representation, evaluation stage, improvement stage, and final output action-value based.

  • Giving only the target name

    A complete algorithm must connect the representation, the iteration process, and the q* result.

    Fix: Include the action-value representation, policy evaluation, policy improvement, repetition, and final result.

  • Inventing a numerical environment

    The source provides none of the particular numerical information needed for a calculation.

    Fix: State the missing information and give a general algorithm blueprint instead.

  • Leaving policy evaluation and improvement disconnected

    The policy-iteration process must show how evaluation produces action values and how improvement uses them to update the policy.

    Fix: Present the stages as a cycle: represent, evaluate, improve, repeat, and report q*.

Practice: Audit an Algorithm

MEDIUM

A proposed answer says: Start with a policy, evaluate its state values, improve the policy, and repeat until the result is v*. The exercise, however, asks for the action-value version and the target q*. Rewrite the answer as an action-value algorithm.

Hints
  • Replace the state-value representation with an action-value representation.
  • Name what policy evaluation produces in the adapted procedure.
  • Explain what policy improvement uses to form the next policy.
  • End by identifying the result as q*, while avoiding invented numerical details.

A strong revision should mention q-values throughout the process, not only in the final line.

Final Checklist

  1. The action-value target named by the exercise is q*, in analogy with v*.
  2. The central task is to reformulate policy iteration so that action values are the main quantities represented and updated.
  3. A complete algorithm connects the action-value representation, policy evaluation, policy improvement, repeated iteration, and the final q* result.
  4. The supplied material does not support a numerical q* calculation because it omits a particular environment, action set, reward description, initial policy, and earlier v* algorithm.
  5. When numerical inputs are missing, provide a precise algorithmic blueprint rather than inventing values.

Key Takeaways

  • Policy iteration is applicable to action values as well as state values.
  • The action-value exercise targets q*, while the state-value target is v*.
  • The adaptation must use action-value language through representation, evaluation, improvement, and output.
  • No numerical q* result can be computed without the environment and other required problem information.