Concepts / State-Value Functions

State-Value Functions

Optimality is evaluated across all states, not from one state alone.

  • Programming

Why One Starting State Is Not Enough

In reinforcement learning, solving a task means finding a policy that produces substantial reward over the long run. However, a policy is not called optimal merely because it performs well in one situation or from one starting state. Optimality is evaluated across all states.

Comparing Policies State by State

Let vπ(s) represent the state value produced by policy π at state s. Policy π is better than or equal to policy π′ when vπ(s) is greater than or equal to vπ′(s) for every state. In plain language, the first policy must produce an expected return at least as high as the second policy from every state being considered.

comparecomparecomparecomparecomparecomparePolicy Astate valuesState 1vA(1) >= vB(1)State 2vA(2) >= vB(2)Policy Bstate valuesState 3vA(3) >= vB(3)
How can one policy be better than or equal to another when their state values are compared at every state?

A policy that is better in every state

Compare policy A with policy B when A's corresponding state values are at least as large in every state.

Compare the first state: Policy A's value is at least as large as policy B's value in the first state.

Compare the remaining states: The same relationship holds in every other state.

Apply the definition: Because policy A is not worse in any state, policy A is better than or equal to policy B.

Policy A is better than or equal to policy B.

Why the Ordering Is Partial

The comparison rule does not necessarily place every pair of policies in order. One policy may have a higher value in one state, while another policy has a higher value in a different state. When that happens, neither policy is better than or equal to the other under the rule, because each policy is worse somewhere.

valuevaluevaluevaluevaluevaluePolicy Ahigher in State 1State 1A > CState 2C > APolicy Chigher in State 2State 3A = C
What does it look like when one policy is better in some states, another is better in others, and neither is worse in none?

Suppose policy A is better in the first state, policy C is better in the second state, and the two policies tie in the third state. Policy A cannot be declared better than or equal to policy C because A loses in the second state. Policy C cannot be declared better than or equal to policy A because C loses in the first state. The policies therefore cannot be ordered by this rule.

A partial ordering allows some policies to be compared while leaving other pairs incomparable. The state-by-state requirement is what creates this limitation.

Testing Optimality in Every State

A policy is optimal when its expected return is at least as high as the expected return of every other policy, starting from every state. In a finite task, at least one policy is better than or equal to all other policies. Every policy meeting this standard is an optimal policy, denoted by π∗.

evaluateevaluateevaluatecompare with all policiescompare with all policiescompare with all policiesPolicy πcandidate policyState 1expected returnOptimalat least as high as everypolicyState 2expected returnState 3expected return
How does evaluating a policy from every state identify whether it is optimal rather than evaluating it from only one starting state?

The Finite Decision Process

In a finite Markov decision process, the policy is evaluated through the states, available actions, transitions, and rewards that make up the task. The policy specifies a way of behaving, and its state-value function records the expected return associated with following that policy from each state. Optimality means that this behavior produces an expected return at least as high as every alternative policy from every state.

acts fromchooses amongaffectsaffectsproducessupports comparisonStatesPolicyway of behavingState-value functionexpected return by stateOptimal policybest from every stateAvailable actionsTransitionsRewards
How do states, available actions, transitions, rewards, and a policy connect to determine whether the policy is optimal?

When Optimal Policies Tie

Optimality does not require a unique policy. There may be more than one policy that is better than or equal to all other policies. All policies meeting this standard are denoted by π∗.

may choosemay choosesharessharesπ∗1optimal policyAction choicepolicy π∗1v∗one optimal state-valuefunctionπ∗2optimal policyAction choicepolicy π∗2
How can different policies choose different actions while producing the same optimal state-value function for every state?

An optimal policy identifies a way of behaving. The optimal state-value function, written v∗, identifies the best expected return associated with each state. Therefore, multiple optimal policies can exist while sharing exactly the same state-value function.

hashasequalsequalsπ∗1optimal policyvπ∗1state valuesv∗shared optimal valuesπ∗2optimal policyvπ∗2state values
How do the state values of multiple optimal policies relate to one shared optimal value for each state?

Common Reasoning Errors

  • Declaring a policy optimal because it performs best in one state.

    Optimality requires the expected return to be at least as high as every other policy from every state.

    Fix: Compare the policies across all states before making an optimality claim.

  • Assuming that every pair of policies must be orderable.

    Each policy is worse in at least one state, so neither satisfies the better-than-or-equal comparison against the other.

    Fix: Allow the result that the policies are incomparable under the state-by-state rule.

  • Assuming that there can be only one optimal policy.

    The definition permits more than one policy to meet the optimality standard.

    Fix: Distinguish the possibility of multiple optimal policies from the shared optimal state-value function.

  • Treating an optimal policy and the optimal state-value function as the same thing.

    A policy identifies behavior, while v∗ identifies the best expected return associated with each state.

    Fix: Use policy language for behavior and state-value-function language for expected returns by state.

Check Your Understanding

MEDIUM

A finite task has three policies. Policy A has a value at least as high as policy B in every state. Policy C is higher than policy A in one state but lower than policy A in another. Which policy comparison can be established, and what can you conclude about whether A and C are ordered by the comparison rule?

Hints
  • Apply the comparison separately to every state.
  • A policy must not be lower in any state to be better than or equal to another.

What do you think happens?

If two different policies are both optimal, must their state-value functions be different?

  • Yes, because different policies must always have different values
  • No, they can share the same optimal state-value function
  • Only if the task has one state
  • The definition requires exactly one policy and one value function
Reveal answer

Answer: No, they can share the same optimal state-value function.

An optimal policy identifies a way of behaving, while v∗ identifies the best expected return associated with each state. Multiple optimal policies can therefore share exactly the same state-value function.

Essential Takeaways

  1. Policy π is better than or equal to policy π′ when its state value is at least as high in every state.
  2. Because one policy can be better in some states while another is better in others, not every pair of policies can be ordered.
  3. In a finite task, an optimal policy is at least as good as every other policy from every state.
  4. There may be multiple optimal policies, and all of them can share one optimal state-value function, v∗.
  5. A policy describes behavior, while a state-value function describes expected return associated with states.

Key Takeaways

  • State-value comparisons must be made across all states.
  • A policy is better than or equal to another only when it is not worse in any state.
  • This rule creates a partial ordering because some policies are better in different states and therefore cannot be ordered.
  • An optimal policy is at least as good as every other policy from every state.
  • Multiple optimal policies can share the same optimal state-value function, v∗.