Concepts / Deterministic Policies

Deterministic Policies

The value function vπ gives the value of following policy π from state s.

  • Programming

A Better-Policy Search

When searching for a better policy, a different action should not be judged only by how attractive it looks in isolation. The important question is what happens when that action is chosen in state s and the existing policy is followed afterward. The value function gives the baseline needed for this comparison: vπ gives the value of following policy π from state s.

The value function is both a description of the current policy and a reference point for judging a possible change.

followgivesconsiderevaluate with π afterwardbaselineState sPolicy πcurrent behaviorvπ(s)baseline valueAction apossible changePolicy comparisoncandidate versus baseline
How does the value function show the expected return from each state when the existing policy is followed?

Two Behaviors from State s

There are two behaviors to compare. In the first, the existing policy π is followed immediately from state s. Its value is vπ(s). In the second, a proposed action a is selected in state s, and policy π is followed thereafter. The proposed behavior is therefore not just action a by itself. It is the combined behavior of choosing a now and using the existing policy as the continuation.

follow immediatelyevaluatedeviate oncethenevaluateState sPolicy πimmediate choicevπ(s)current valueAction aproposed choicePolicy πcontinuationqπ(s,a)proposed value
How does taking action a in state s and then following policy π compare with following π from s without deviating?

Comparing Two Continuations

Suppose state s has a current value vπ(s) of 10. A proposed action a, followed by policy π thereafter, has value qπ(s,a) of 12.

Identify the baseline: The baseline is vπ(s), because it represents following the existing policy from state s.

Identify the proposed behavior: The proposal is not action a in isolation. It is action a in state s followed by policy π afterward.

Compare the values: The proposed value, 12, is greater than the current-policy value, 10.

Interpret the result: The policy evaluation criterion supports changing the choice in state s because the proposed behavior is better than following the existing policy immediately.

The proposed action is promising under the policy evaluation criterion.

The Policy Evaluation Criterion

The policy evaluation criterion asks whether the proposed behavior has greater value than the current baseline. In the usual notation, qπ(s,a) represents the value of selecting action a in state s and then following π, while vπ(s) represents the value of following π from s. A proposed change is supported when qπ(s,a) is greater than vπ(s). If the proposed behavior does not compare favorably with vπ(s), the criterion does not support the change.

greater than currentcomparison baselinenot greaterbaseline remains preferableqπ(s,a)action a, then πvπ(s)π from sPolicy changesupported when candidate isgreaterExisting policychange not supported
How can we compare qπ(s,a) with vπ(s) to decide whether changing the action at state s is worthwhile?

When a Change Is Not Supported

Suppose the current value vπ(s) is 10, while choosing action a in s and then following π has value qπ(s,a) of 8.

Set the comparison: Compare the value of the proposed behavior with the value of following π immediately.

Read the result: The proposed value, 8, is less than the baseline value, 10.

Apply the criterion: Because the proposed behavior has no greater value, the policy evaluation criterion does not support changing the action.

The existing policy remains the supported choice for this comparison.

From a Promising Choice to Improvement

The Policy Improvement Theorem connects the one-state comparison to a new deterministic policy. If selecting the proposed action once in state s and then following π is better than following π from the beginning, the reasoning is that selecting that action every time state s is encountered should be better still. The new policy applies the proposed choice consistently. For deterministic policies, this is identified as a policy-improvement result.

usesselectsapplies consistentlyState sbeforeState safterPolicy πcurrent choiceAction achosen consistentlyNew policydeterministic choice
How does choosing actions that are at least as good as the current policy guarantee that the new deterministic policy is no worse?

The theorem is about a policy change, not merely an isolated action. The proposed action becomes part of a new policy that applies the choice consistently whenever the state is encountered.

What do you think happens?

If taking action a once in state s and then following π is better than following π immediately, what does the Policy Improvement Theorem suggest about choosing a in state s consistently?

  • The consistent choice is a policy-improvement result
  • The action must be rejected because it was tested only once
  • The current policy is automatically better
Reveal answer

Answer: The consistent choice is a policy-improvement result.

The theorem uses the one-step comparison as the basis for concluding that applying the proposed choice consistently in the new deterministic policy is better or no worse than the current policy.

Mistakes in Policy Comparison

  • Judging an action in isolation

    The evaluation concerns action a followed by policy π, not action a separated from its continuation.

    Fix: Compare the complete proposed behavior with vπ(s).

  • Using the wrong baseline

    The value function vπ(s) is the baseline for the current policy.

    Fix: Start the comparison with the value of following π immediately from s.

  • Treating a promising one-step change as unrelated to the new policy

    The Policy Improvement Theorem explains why applying the proposed choice consistently produces a policy-improvement result for deterministic policies.

    Fix: Connect the one-step comparison to the new policy that consistently selects the proposed action.

For every proposed action, write down the two sides of the comparison explicitly: the current-policy value vπ(s), and the value of taking action a once in s before continuing with π. This keeps the continuation visible and prevents an isolated-action judgment.

Apply the Criterion

MEDIUM

A state s has current-policy value vπ(s) equal to 15. A proposed action a followed by policy π has value qπ(s,a) equal to 15. Decide whether the policy evaluation criterion supports the change. Then explain what the Policy Improvement Theorem says if the new deterministic policy applies the proposed choice consistently.

Hints
  • Compare the proposed value with the current-policy value.
  • The criterion supports a change when the proposed behavior has greater value.
  • Remember that the theorem concerns the consistently applied choice in the new deterministic policy.

In this case, the proposed behavior is equal in value to the current baseline, not greater. Therefore the strict policy evaluation criterion described here does not support the change as an improvement. The theorem's relevant condition is that the proposed behavior be at least as good when considering the new deterministic policy; the source emphasizes that a greater proposed value supports the change and that the new policy is no worse when its choices are at least as good as the current policy.

Key Takeaways

  1. vπ(s) is the value of following the existing policy π from state s.
  2. A candidate action must be evaluated together with the continuation of following π afterward.
  3. The policy evaluation criterion compares the candidate behavior with vπ(s); a greater candidate value supports the change.
  4. The value function serves as the baseline for searching for a better policy.
  5. The Policy Improvement Theorem explains why a promising one-step choice can support a new deterministic policy that applies that choice consistently.

Key Takeaways

  • Use vπ(s) as the baseline for the current policy.
  • Compare following π immediately with taking action a once and then following π.
  • A proposed change is supported when the proposed behavior has greater value than vπ(s).
  • The Policy Improvement Theorem links this comparison to consistently applying the proposed choice in a new deterministic policy.