Reinforcement Learning Optimality
Value functions evaluate expected returns under a specified policy.
From Evaluation to Choice
Reinforcement learning separates two questions that are easy to confuse. First, how well does a particular policy perform? Second, what is the largest expected return that any policy can achieve? The first question is policy evaluation. The second is optimality. Understanding the difference is the foundation for understanding optimal value functions and optimal policies.
A value function is not a permanent score attached to a state. Its meaning depends on the policy being followed.
What a Policy Evaluates
Start by fixing a policy. A policy specifies the agent's behavior, and its value functions evaluate the expected return from that behavior. A state value evaluates the expected return when the agent starts from a state and follows the chosen policy from that point onward. A state-action value evaluates the expected return when evaluation begins with a particular state-action pair, while the policy still governs the agent's behavior.
Evaluating Two Starting Points
Imagine an agent following one fixed policy. Compare asking for the value of state S with asking for the value of the pair consisting of state S and action A.
State evaluation: The state value asks what expected return follows when the agent is at S and uses the fixed policy from there.
State-action evaluation: The state-action value begins with the particular pair S and A, then continues under the policy's behavior.
Meaning of both results: Both results are expected-return assessments under the same policy. They differ in where evaluation begins.
A policy's value functions evaluate future prospects from either a state or a specified state-action pair; neither is an unconditional score for the state.
Ordinary and Optimal Values
An ordinary value function evaluates one specified policy. It answers, “How much return should be expected if this policy is followed?” An optimal value function changes the question to, “What is the largest expected return achievable by any policy?” The optimal value therefore represents the best achievable performance, rather than the performance of one already-selected policy.
| Question | Value function | Meaning |
|---|---|---|
| How well does this policy perform? | Value function for a specified policy | Expected return under that policy |
| What is the best return any policy can achieve? | Optimal value function | Largest expected return achievable by any policy |
The defining distinction between ordinary and optimal value functions.
From Values to Policies
Optimal value functions and optimal policies are connected. An optimal policy is a policy whose value functions are optimal. Once the optimal values are known, they can be used to determine an optimal policy. In practice, a policy that is greedy with respect to the optimal value functions must be optimal.
Selecting an Action from Optimal Values
Suppose a state has several available actions, and the optimal value assessment identifies one action as producing the largest achievable expected return from that state.
Inspect the alternatives: The optimal assessment considers what can be achieved through the available policy choices.
Identify the largest value: The action associated with the largest achievable expected return is the greedy choice with respect to the optimal values.
Construct the policy: Selecting such actions across states gives a policy whose value functions are optimal.
Optimal values can be used to determine an optimal policy, and a greedy policy with respect to those values must be optimal.
The optimal values are unique for a given MDP, but the optimal policy does not have to be unique. Multiple policies may achieve the same optimal values.
Bellman Optimality Conditions
Bellman optimality equations provide consistency conditions for optimal value functions. They express the requirement that the values assigned as optimal must agree with the best achievable choices throughout the decision process. In principle, solving these equations yields the optimal value functions. An optimal policy can then be determined from those values.
- Begin with the requirement that the values represent the largest expected returns achievable by any policy.
- Use the Bellman optimality equations as consistency conditions for those values.
- Solve the conditions in principle to obtain the optimal value functions.
- Use the resulting optimal values to determine an optimal policy.
Common Reasoning Errors
Treating a value as a permanent score attached to a state.
A value function's meaning depends on the policy being followed.
Fix:
Always state which policy is being evaluated before interpreting a value.Calling the value of one policy the optimal value.
A policy value evaluates one specified policy, while an optimal value is the largest expected return achievable by any policy.
Fix:
Separate the question of policy performance from the question of best achievable performance.Assuming that unique optimal values imply one unique optimal policy.
For a given MDP, optimal values are unique, but multiple policies may share them.
Fix:
Treat the optimal value as fixed while allowing more than one policy to attain it.Using Bellman optimality equations as though they only evaluate a chosen policy.
They are consistency conditions for optimal value functions.
Fix:
Understand their role as conditions that can be solved for optimal values, followed by policy determination.
Check Your Understanding
A learner says: “The value of state S is 8, so state S always has value 8.” Explain why this statement is incomplete. Then contrast the meaning of the value of S under one specified policy with the meaning of the optimal value of S.
Hints
- Ask what behavior the ordinary value function assumes.
- Ask whether the optimal value considers one policy or all policies.
- Remember that optimal values can be shared by multiple optimal policies.
What do you think happens?
If two different policies achieve the same optimal values for a given MDP, must one of them be non-optimal?
Reveal answer
Answer: No, multiple policies may attain the same optimal values.
The optimal values for a given MDP are unique, but the policies that achieve them do not have to be unique.
Key Takeaways
- A policy's value functions assign expected returns to states and to state-action pairs under that policy.
- Ordinary value functions evaluate one specified policy.
- Optimal value functions select the largest expected return achievable by any policy.
- Optimal values can determine optimal policies, and several optimal policies may share the same optimal values.
- Bellman optimality equations provide consistency conditions that can be solved for optimal values before determining an optimal policy.
Key Takeaways
- Value functions are policy-dependent expected-return evaluations.
- State values begin from a state, while state-action values begin from a specified state-action pair.
- Optimal value functions represent the largest expected return achievable by any policy.
- An optimal policy can be determined from optimal values, and more than one policy may attain the same optimal values.
- Bellman optimality equations state consistency conditions for finding optimal value functions.