Policy Improvement
Policy evaluation measures a fixed policy rather than changing it.
From Evaluation to Improvement
Policy evaluation and policy improvement answer different questions. Evaluation asks: how good is this fixed policy? Improvement asks: can we change the policy without making its expected results worse? Iterative policy evaluation approaches the value of a fixed policy through successive approximations. It does not change the policy while measuring it.
Keep the policy fixed while evaluating it. Change the policy only when making an improvement decision.
Repeated Estimates in a Gridworld
Imagine a gridworld with an equiprobable random policy: whenever several actions are available, the policy selects them with equal probability. Iterative evaluation begins with an approximation to the values, then repeatedly updates those estimates according to the fixed policy. Each round uses the current estimates to produce a better approximation of what states are worth under that policy.
Setting Up a Gridworld Evaluation
A gridworld uses an equiprobable random policy. A new state is introduced, and its value must be evaluated under that unchanged policy. What should be held fixed, and what should be recomputed?
Hold the policy fixed: Keep the equal-probability action selection rule unchanged. The task is policy evaluation, not policy improvement.
Reflect the gridworld change: Include the new state and its transitions in the evaluation problem before recomputing values.
Iterate the estimates: Use successive approximations to determine the value of the changed state and the other affected states under the fixed policy.
The gridworld change must be represented first, and then the value function for the unchanged policy must be recomputed.
State Values and Action Values
The state-value function vπ(s) describes the expected return from starting in state s and then following policy π. The action-value function qπ(s, a) describes the expected return from starting in state s, taking the specified action a first, and then following policy π. The difference is the first decision being described: vπ begins with the policy’s behavior, while qπ fixes the first action before policy behavior continues.
| Quantity | What is fixed first? | What follows afterward? |
|---|---|---|
| vπ(s) | Starting state s | Behavior selected by policy π |
| qπ(s, a) | Starting state s and first action a | Behavior selected by policy π |
Reading the Two Value Functions
At state s, compare the meanings of vπ(s) and qπ(s, a).
Read vπ(s): Interpret it as the expected return obtained by starting at s and behaving according to π.
Read qπ(s, a): Interpret it as the expected return obtained by starting at s, forcing action a as the first action, and following π afterward.
Use the distinction: Policy improvement needs qπ because it compares a candidate first action with the original policy’s state value vπ(s).
vπ(s) evaluates the original policy from a state; qπ(s, a) evaluates a specified first action followed by that policy.
When Evaluation Cannot Settle
Undiscounted episodic tasks are intended to terminate, but a particular policy may fail to reach a terminal outcome. In the grid problem, a policy can create a two-state loop in which the agent moves back and forth forever. For some policies and starting states, vπ(s) may therefore be negative infinity. The iterative policy-evaluation procedure may not terminate in that situation.
Assuming that an episodic task guarantees termination under every policy.
A particular policy can keep the agent in a two-state loop forever.
Fix:
Check whether the policy itself guarantees reaching termination before assuming iterative evaluation will settle.Treating nontermination as only a value-calculation issue.
The evaluation computation itself may fail to terminate, and some policy values may be negative infinity.
Fix:
Modify the evaluation procedure with a practical stopping safeguard.
The Improvement Guarantee
Begin with an original deterministic policy π and a changed deterministic policy π′. At each state s, the original policy’s expected return is vπ(s). The action selected by the changed policy is π′(s), and its expected return is measured as qπ(s, π′(s)): take the new policy’s selected action once, then follow the original policy π.
qπ(s, π′(s)) ≥ vπ(s) for every state s
The theorem connects a local comparison with the value of following the entire changed policy. Its proof idea repeatedly expands the qπ side and reapplies the comparison condition until the reasoning reaches vπ′(s). Thus, the guarantee is not limited to the first changed action: it concerns the expected return under the whole new policy.
Testing a Candidate Policy Change
A candidate deterministic policy π′ is compared with an original deterministic policy π. At every state, qπ(s, π′(s)) is equal to or greater than vπ(s), and at one state it is strictly greater. What does the theorem guarantee?
Check coverage: The comparison is made for every state, so the theorem’s state-by-state condition is satisfied.
Check strictness: At least one state has a strict improvement in the candidate action’s expected return.
Apply the conclusion: The changed policy is guaranteed to be at least as good as the original policy and strictly better at least at one state.
π′ is guaranteed not to reduce policy value and must improve expected return at at least one state.
Evaluation Versus Improvement
| Activity | What stays fixed? | What is computed or compared? |
|---|---|---|
| Policy evaluation | The policy π | The value function for that policy |
| Policy improvement | The original policy’s values used for comparison | Candidate actions qπ(s, π′(s)) versus vπ(s) |
| Theorem conclusion | The inequality condition at every state | Whether following π′ is guaranteed to be no worse or better |
Comparing qπ(s, π′(s)) with the wrong baseline.
The theorem compares the new action’s expected return with the original policy’s state value.
Fix:
Use qπ(s, π′(s)) ≥ vπ(s) as the state-by-state test.Checking the inequality at only the states where the policy changes.
The theorem requires the condition for every state.
Fix:
Verify the inequality across the complete state set.Assuming equality everywhere proves strict improvement.
Equality everywhere guarantees no worse, not strictly better.
Fix:
Look for a strict inequality at at least one state before claiming strict improvement.Treating policy evaluation as if it changes the policy.
Evaluation measures a fixed policy.
Fix:
Hold π fixed during evaluation; use a separate improvement step to construct π′.
Apply the Statewise Test
For a proposed deterministic policy π′, decide whether the policy improvement theorem applies. Check whether qπ(s, π′(s)) is at least vπ(s) for every state. Then determine whether the conclusion is merely no worse or strictly better somewhere.
Hints
- The condition must hold at every state.
- Equality at all states gives a no-worse guarantee.
- A strict inequality at one state gives strict improvement at at least one state.
What do you think happens?
Suppose the candidate action is better at one state, equal at all other states, and never worse anywhere. Is the changed policy guaranteed to be at least as good as the original policy?
Reveal answer
Answer: Yes, and it is guaranteed to be strictly better at at least one state.
The theorem requires a non-strict comparison at every state and a strict comparison at at least one state. Both conditions are satisfied.
Key Takeaways
- Iterative policy evaluation uses successive approximations to measure a fixed policy.
- vπ(s) evaluates starting in a state and then following π; qπ(s, a) evaluates taking a specified first action and then following π.
- A gridworld transition change must be included before values are recomputed.
- An undiscounted policy can loop forever, so evaluation may need an explicit stopping safeguard.
- If qπ(s, π′(s)) ≥ vπ(s) for every state, the changed deterministic policy π′ is guaranteed to be no worse; if the inequality is strict somewhere, π′ is guaranteed to be strictly better at at least one state.
Key Takeaways
- Policy evaluation measures a fixed policy rather than changing it.
- State values describe states under a policy, while action values describe a specified state-action pair followed by that policy.
- Undiscounted policies can create nonterminating loops, so practical evaluation needs a stopping safeguard.
- The policy improvement theorem guarantees that a deterministic policy change is no worse when its selected action satisfies qπ(s, π′(s)) ≥ vπ(s) at every state.
- A strict inequality at one state guarantees strict improvement at at least one state.