Value Function for a Policy
The value function vπ gives the value of following policy π from state s.
Why the Current Policy Needs a Baseline
When searching for a better policy, it is not enough to notice that a different action looks attractive by itself. The important question is whether that action produces better behavior when it is followed by the existing policy afterward. The value function gives the baseline needed for this comparison.
The value function vπ gives the value of following policy π from state s.
The notation vπ(s) represents the current policy's value when starting in state s and following π. This value is more than a description of what the current policy does. It is the reference point against which a possible policy change is judged.
The One-Step Alternative
To evaluate a proposed change in state s, compare two complete behavior patterns. The baseline is to follow π immediately from s. The alternative is to select a proposed action a in s and then follow π thereafter. The action is therefore evaluated together with its continuation, not in isolation.
Comparing Two Behaviors from One State
In state s, decide whether replacing the current choice with proposed action a is supported by the policy evaluation criterion.
Set the baseline: Use vπ(s), the value of starting in s and following the existing policy π.
Evaluate the alternative: Consider selecting a in s and then following π thereafter. This is the behavior represented by qπ(s, a).
Compare the values: If qπ(s, a) is greater than vπ(s), the proposed behavior has greater value than the baseline.
Interpret the result: The policy evaluation criterion supports changing the choice in state s when the proposed action followed by π compares favorably with following π immediately.
The decision depends on the comparison between qπ(s, a) and vπ(s), not on whether action a looks attractive in isolation.
The State-by-State Evaluation Test
The policy evaluation criterion is applied state by state. For a state s and proposed action a, compare qπ(s, a) with vπ(s). The first quantity represents selecting a in s and then following π. The second represents following π from the beginning. If the proposed behavior has greater value, the criterion supports the change at that state.
| Behavior | Meaning | Comparison role |
|---|---|---|
| Following π immediately | Start in s and follow the existing policy | Baseline vπ(s) |
| Selecting a, then following π | Take the proposed action in s and use π afterward | Alternative qπ(s, a) |
From Local Choices to a New Policy
A new deterministic policy π′ chooses one action for each state. The Policy Improvement Theorem connects the one-step comparison to this new policy. If qπ(s, π′(s)) is at least vπ(s) for every state, then the new deterministic policy π′ is no worse than π.
The theorem's reasoning is that if taking the proposed action once and then following π is at least as good as following π immediately in every state, then choosing that action consistently whenever the state is encountered produces a deterministic policy that is no worse than the original policy. The theorem therefore turns state-by-state action comparisons into a policy-level conclusion.
A Generated State-by-State Scenario
Imagine a policy π and a state s. Suppose the current policy's value is represented by vπ(s) = 8. A proposed action a, followed by π afterward, has value qπ(s, a) = 10. Because 10 is greater than 8, the policy evaluation criterion supports changing the choice in state s. If instead qπ(s, a) were 6, the comparison would not support that change.
If the proposed behavior does not compare favorably with vπ(s), the policy evaluation criterion does not support the change at that state. The action should not be accepted merely because it appears attractive before its continuation under π is considered.
Mistakes in Policy Comparison
Judging a proposed action in isolation
The relevant alternative is selecting a in s and then following π thereafter.
Fix:
Compare qπ(s, a) with vπ(s).Forgetting the baseline
A policy change is evaluated relative to the current value vπ(s).
Fix:
Use vπ(s) as the baseline for every state-level comparison.Assuming one favorable state proves the whole policy is better
The Policy Improvement Theorem uses the condition qπ(s, π′(s)) is at least vπ(s) for every state.
Fix:
Check the comparison state by state before making the policy-level conclusion.Confusing a one-step test with a complete replacement of the continuation
The defined proposed behavior continues with the existing policy π.
Fix:
Keep π as the continuation when evaluating the proposed first action.
When evaluating a policy change, write down the two behaviors before comparing them: follow π from s, or select the proposed action in s and then follow π. This prevents the continuation from disappearing from the analysis.
Practice: Apply the Criterion
For a state s, the existing policy has value vπ(s) = 12. A proposed action followed by π has value qπ(s, a) = 12. Decide whether the policy evaluation criterion supports the change at this state. Then explain whether this single comparison is enough to conclude that a new deterministic policy is no worse than π across all states.
Hints
- Compare the two values using the criterion.
- Remember that the proposed behavior includes both the action and the continuation with π.
- For the theorem-level conclusion, consider whether the condition has been checked for every state.
What do you think happens?
Does equality between qπ(s, a) and vπ(s) support a change at this state, and does one state's equality establish that the new deterministic policy is no worse everywhere?
Reveal answer
Answer: The state-level comparison is not worse, but the policy-level conclusion still requires the condition for every state
The theorem condition uses qπ(s, π′(s)) at least vπ(s) for every state. Equality satisfies the no-worse comparison at one state, but one state alone is not enough for the policy-level result.
Key Takeaways
- The value function vπ gives the value of following policy π from state s.
- It serves as the baseline for judging a proposed policy change.
- The meaningful alternative is to take a proposed action in s and then follow π, not to judge the action alone.
- The policy evaluation criterion supports a change when the proposed behavior compares favorably with vπ(s).
- The Policy Improvement Theorem states that a deterministic policy π′ is no worse than π when qπ(s, π′(s)) is at least vπ(s) for every state.
Key Takeaways
- Use vπ(s) as the current policy's baseline value from state s.
- Compare following π immediately with taking a proposed action first and then following π.
- A favorable qπ(s, a) versus vπ(s) comparison supports a change at that state.
- For a deterministic policy π′ to be no worse than π under the Policy Improvement Theorem, qπ(s, π′(s)) must be at least vπ(s) for every state.