Concepts / Value Function for a Policy

Value Function for a Policy

The value function vπ gives the value of following policy π from state s.

  • Programming

Why the Current Policy Needs a Baseline

When searching for a better policy, it is not enough to notice that a different action looks attractive by itself. The important question is whether that action produces better behavior when it is followed by the existing policy afterward. The value function gives the baseline needed for this comparison.

The value function vπ gives the value of following policy π from state s.

The notation vπ(s) represents the current policy's value when starting in state s and following π. This value is more than a description of what the current policy does. It is the reference point against which a possible policy change is judged.

start inevaluateState sPolicy πfollow πvπ(s)value of following π
What future behavior is represented by vπ(s) when starting in state s?

The One-Step Alternative

To evaluate a proposed change in state s, compare two complete behavior patterns. The baseline is to follow π immediately from s. The alternative is to select a proposed action a in s and then follow π thereafter. The action is therefore evaluated together with its continuation, not in isolation.

selectthen followfollow immediatelyState sState sAction acandidate choicePolicy πimmediate continuationPolicy πcontinuation
What is the difference between taking a candidate action first and following π immediately from the current state?

Comparing Two Behaviors from One State

In state s, decide whether replacing the current choice with proposed action a is supported by the policy evaluation criterion.

Set the baseline: Use vπ(s), the value of starting in s and following the existing policy π.

Evaluate the alternative: Consider selecting a in s and then following π thereafter. This is the behavior represented by qπ(s, a).

Compare the values: If qπ(s, a) is greater than vπ(s), the proposed behavior has greater value than the baseline.

Interpret the result: The policy evaluation criterion supports changing the choice in state s when the proposed action followed by π compares favorably with following π immediately.

The decision depends on the comparison between qπ(s, a) and vπ(s), not on whether action a looks attractive in isolation.

The State-by-State Evaluation Test

The policy evaluation criterion is applied state by state. For a state s and proposed action a, compare qπ(s, a) with vπ(s). The first quantity represents selecting a in s and then following π. The second represents following π from the beginning. If the proposed behavior has greater value, the criterion supports the change at that state.

qπ(s, a) greater than vπ(s)does not compare favorablyCompare valuesqπ(s, a) and vπ(s)Policy changeproposed behavior hasgreater valueExisting policycriterion does not supportchange
How does comparing qπ(s, a) with vπ(s) determine whether a policy change is promising?
BehaviorMeaningComparison role
Following π immediatelyStart in s and follow the existing policyBaseline vπ(s)
Selecting a, then following πTake the proposed action in s and use π afterwardAlternative qπ(s, a)

From Local Choices to a New Policy

A new deterministic policy π′ chooses one action for each state. The Policy Improvement Theorem connects the one-step comparison to this new policy. If qπ(s, π′(s)) is at least vπ(s) for every state, then the new deterministic policy π′ is no worse than π.

baseline for comparisonat least vπ(s) for every statevπ(s)existing policy baselineqπ(s, π′(s))new policy's selectedactionπ′no worse than π
How does showing qπ(s, π′(s)) is at least vπ(s) for every state establish that π′ is no worse than π?

The theorem's reasoning is that if taking the proposed action once and then following π is at least as good as following π immediately in every state, then choosing that action consistently whenever the state is encountered produces a deterministic policy that is no worse than the original policy. The theorem therefore turns state-by-state action comparisons into a policy-level conclusion.

inspectchooseState sAction valuecompare with vπ(s)π′(s)chosen action
How does a new deterministic policy choose an action for each state by comparing available action values?

A Generated State-by-State Scenario

Imagine a policy π and a state s. Suppose the current policy's value is represented by vπ(s) = 8. A proposed action a, followed by π afterward, has value qπ(s, a) = 10. Because 10 is greater than 8, the policy evaluation criterion supports changing the choice in state s. If instead qπ(s, a) were 6, the comparison would not support that change.

comparegreater than baselinevπ(s) = 8follow π immediatelyqπ(s, a) = 10take a, then follow πChange supportedalternative is greater
What changes when the proposed action followed by π has greater value than following π immediately?

If the proposed behavior does not compare favorably with vπ(s), the policy evaluation criterion does not support the change at that state. The action should not be accepted merely because it appears attractive before its continuation under π is considered.

Mistakes in Policy Comparison

  • Judging a proposed action in isolation

    The relevant alternative is selecting a in s and then following π thereafter.

    Fix: Compare qπ(s, a) with vπ(s).

  • Forgetting the baseline

    A policy change is evaluated relative to the current value vπ(s).

    Fix: Use vπ(s) as the baseline for every state-level comparison.

  • Assuming one favorable state proves the whole policy is better

    The Policy Improvement Theorem uses the condition qπ(s, π′(s)) is at least vπ(s) for every state.

    Fix: Check the comparison state by state before making the policy-level conclusion.

  • Confusing a one-step test with a complete replacement of the continuation

    The defined proposed behavior continues with the existing policy π.

    Fix: Keep π as the continuation when evaluating the proposed first action.

When evaluating a policy change, write down the two behaviors before comparing them: follow π from s, or select the proposed action in s and then follow π. This prevents the continuation from disappearing from the analysis.

Practice: Apply the Criterion

MEDIUM

For a state s, the existing policy has value vπ(s) = 12. A proposed action followed by π has value qπ(s, a) = 12. Decide whether the policy evaluation criterion supports the change at this state. Then explain whether this single comparison is enough to conclude that a new deterministic policy is no worse than π across all states.

Hints
  • Compare the two values using the criterion.
  • Remember that the proposed behavior includes both the action and the continuation with π.
  • For the theorem-level conclusion, consider whether the condition has been checked for every state.

What do you think happens?

Does equality between qπ(s, a) and vπ(s) support a change at this state, and does one state's equality establish that the new deterministic policy is no worse everywhere?

  • Yes at the state, and yes for every state
  • No at the state, and no for every state
  • The state-level comparison is not worse, but the policy-level conclusion still requires the condition for every state
Reveal answer

Answer: The state-level comparison is not worse, but the policy-level conclusion still requires the condition for every state

The theorem condition uses qπ(s, π′(s)) at least vπ(s) for every state. Equality satisfies the no-worse comparison at one state, but one state alone is not enough for the policy-level result.

Key Takeaways

  1. The value function vπ gives the value of following policy π from state s.
  2. It serves as the baseline for judging a proposed policy change.
  3. The meaningful alternative is to take a proposed action in s and then follow π, not to judge the action alone.
  4. The policy evaluation criterion supports a change when the proposed behavior compares favorably with vπ(s).
  5. The Policy Improvement Theorem states that a deterministic policy π′ is no worse than π when qπ(s, π′(s)) is at least vπ(s) for every state.

Key Takeaways

  • Use vπ(s) as the current policy's baseline value from state s.
  • Compare following π immediately with taking a proposed action first and then following π.
  • A favorable qπ(s, a) versus vπ(s) comparison supports a change at that state.
  • For a deterministic policy π′ to be no worse than π under the Policy Improvement Theorem, qπ(s, π′(s)) must be at least vπ(s) for every state.