State-Value Functions
Optimality is evaluated across all states, not from one state alone.
Why One Starting State Is Not Enough
In reinforcement learning, solving a task means finding a policy that produces substantial reward over the long run. However, a policy is not called optimal merely because it performs well in one situation or from one starting state. Optimality is evaluated across all states.
Comparing Policies State by State
Let vπ(s) represent the state value produced by policy π at state s. Policy π is better than or equal to policy π′ when vπ(s) is greater than or equal to vπ′(s) for every state. In plain language, the first policy must produce an expected return at least as high as the second policy from every state being considered.
A policy that is better in every state
Compare policy A with policy B when A's corresponding state values are at least as large in every state.
Compare the first state: Policy A's value is at least as large as policy B's value in the first state.
Compare the remaining states: The same relationship holds in every other state.
Apply the definition: Because policy A is not worse in any state, policy A is better than or equal to policy B.
Policy A is better than or equal to policy B.
Why the Ordering Is Partial
The comparison rule does not necessarily place every pair of policies in order. One policy may have a higher value in one state, while another policy has a higher value in a different state. When that happens, neither policy is better than or equal to the other under the rule, because each policy is worse somewhere.
Suppose policy A is better in the first state, policy C is better in the second state, and the two policies tie in the third state. Policy A cannot be declared better than or equal to policy C because A loses in the second state. Policy C cannot be declared better than or equal to policy A because C loses in the first state. The policies therefore cannot be ordered by this rule.
A partial ordering allows some policies to be compared while leaving other pairs incomparable. The state-by-state requirement is what creates this limitation.
Testing Optimality in Every State
A policy is optimal when its expected return is at least as high as the expected return of every other policy, starting from every state. In a finite task, at least one policy is better than or equal to all other policies. Every policy meeting this standard is an optimal policy, denoted by π∗.
The Finite Decision Process
In a finite Markov decision process, the policy is evaluated through the states, available actions, transitions, and rewards that make up the task. The policy specifies a way of behaving, and its state-value function records the expected return associated with following that policy from each state. Optimality means that this behavior produces an expected return at least as high as every alternative policy from every state.
When Optimal Policies Tie
Optimality does not require a unique policy. There may be more than one policy that is better than or equal to all other policies. All policies meeting this standard are denoted by π∗.
An optimal policy identifies a way of behaving. The optimal state-value function, written v∗, identifies the best expected return associated with each state. Therefore, multiple optimal policies can exist while sharing exactly the same state-value function.
Common Reasoning Errors
Declaring a policy optimal because it performs best in one state.
Optimality requires the expected return to be at least as high as every other policy from every state.
Fix:
Compare the policies across all states before making an optimality claim.Assuming that every pair of policies must be orderable.
Each policy is worse in at least one state, so neither satisfies the better-than-or-equal comparison against the other.
Fix:
Allow the result that the policies are incomparable under the state-by-state rule.Assuming that there can be only one optimal policy.
The definition permits more than one policy to meet the optimality standard.
Fix:
Distinguish the possibility of multiple optimal policies from the shared optimal state-value function.Treating an optimal policy and the optimal state-value function as the same thing.
A policy identifies behavior, while v∗ identifies the best expected return associated with each state.
Fix:
Use policy language for behavior and state-value-function language for expected returns by state.
Check Your Understanding
A finite task has three policies. Policy A has a value at least as high as policy B in every state. Policy C is higher than policy A in one state but lower than policy A in another. Which policy comparison can be established, and what can you conclude about whether A and C are ordered by the comparison rule?
Hints
- Apply the comparison separately to every state.
- A policy must not be lower in any state to be better than or equal to another.
What do you think happens?
If two different policies are both optimal, must their state-value functions be different?
Reveal answer
Answer: No, they can share the same optimal state-value function.
An optimal policy identifies a way of behaving, while v∗ identifies the best expected return associated with each state. Multiple optimal policies can therefore share exactly the same state-value function.
Essential Takeaways
- Policy π is better than or equal to policy π′ when its state value is at least as high in every state.
- Because one policy can be better in some states while another is better in others, not every pair of policies can be ordered.
- In a finite task, an optimal policy is at least as good as every other policy from every state.
- There may be multiple optimal policies, and all of them can share one optimal state-value function, v∗.
- A policy describes behavior, while a state-value function describes expected return associated with states.
Key Takeaways
- State-value comparisons must be made across all states.
- A policy is better than or equal to another only when it is not worse in any state.
- This rule creates a partial ordering because some policies are better in different states and therefore cannot be ordered.
- An optimal policy is at least as good as every other policy from every state.
- Multiple optimal policies can share the same optimal state-value function, v∗.