Concepts / Reinforcement Learning Optimality

Reinforcement Learning Optimality

Value functions evaluate expected returns under a specified policy.

  • Programming

From Evaluation to Choice

Reinforcement learning separates two questions that are easy to confuse. First, how well does a particular policy perform? Second, what is the largest expected return that any policy can achieve? The first question is policy evaluation. The second is optimality. Understanding the difference is the foundation for understanding optimal value functions and optimal policies.

A value function is not a permanent score attached to a state. Its meaning depends on the policy being followed.

What a Policy Evaluates

Start by fixing a policy. A policy specifies the agent's behavior, and its value functions evaluate the expected return from that behavior. A state value evaluates the expected return when the agent starts from a state and follows the chosen policy from that point onward. A state-action value evaluates the expected return when evaluation begins with a particular state-action pair, while the policy still governs the agent's behavior.

governs behavior fromgoverns behavior afterevaluatesevaluatesChosen policyStateexpected returnState valueunder the policyState-action pairexpected returnState-action valueunder the policy
What does each value function evaluate when a policy is fixed?

Evaluating Two Starting Points

Imagine an agent following one fixed policy. Compare asking for the value of state S with asking for the value of the pair consisting of state S and action A.

State evaluation: The state value asks what expected return follows when the agent is at S and uses the fixed policy from there.

State-action evaluation: The state-action value begins with the particular pair S and A, then continues under the policy's behavior.

Meaning of both results: Both results are expected-return assessments under the same policy. They differ in where evaluation begins.

A policy's value functions evaluate future prospects from either a state or a specified state-action pair; neither is an unconditional score for the state.

Ordinary and Optimal Values

An ordinary value function evaluates one specified policy. It answers, “How much return should be expected if this policy is followed?” An optimal value function changes the question to, “What is the largest expected return achievable by any policy?” The optimal value therefore represents the best achievable performance, rather than the performance of one already-selected policy.

evaluatesselectsOne specifiedpolicyexpected returnPolicy valueperformance of that policyAny policylargest achievable returnOptimal valuebest achievable performance
What changes when evaluation moves from one specified policy to the best policy available?
QuestionValue functionMeaning
How well does this policy perform?Value function for a specified policyExpected return under that policy
What is the best return any policy can achieve?Optimal value functionLargest expected return achievable by any policy

The defining distinction between ordinary and optimal value functions.

From Values to Policies

Optimal value functions and optimal policies are connected. An optimal policy is a policy whose value functions are optimal. Once the optimal values are known, they can be used to determine an optimal policy. In practice, a policy that is greedy with respect to the optimal value functions must be optimal.

evaluate possibilitiesguideforms behaviorStateOptimal valuesbest achievable prospectsAction choicegreedy with respect tooptimal valuesOptimal policypolicy with optimal valuefunctions
How does an optimal value function guide the action selected by an optimal policy?

Selecting an Action from Optimal Values

Suppose a state has several available actions, and the optimal value assessment identifies one action as producing the largest achievable expected return from that state.

Inspect the alternatives: The optimal assessment considers what can be achieved through the available policy choices.

Identify the largest value: The action associated with the largest achievable expected return is the greedy choice with respect to the optimal values.

Construct the policy: Selecting such actions across states gives a policy whose value functions are optimal.

Optimal values can be used to determine an optimal policy, and a greedy policy with respect to those values must be optimal.

The optimal values are unique for a given MDP, but the optimal policy does not have to be unique. Multiple policies may achieve the same optimal values.

Bellman Optimality Conditions

Bellman optimality equations provide consistency conditions for optimal value functions. They express the requirement that the values assigned as optimal must agree with the best achievable choices throughout the decision process. In principle, solving these equations yields the optimal value functions. An optimal policy can then be determined from those values.

considerevaluateenforceextractStateAvailable actionspossible choicesReturn comparisonidentify the largestachievable returnOptimal valuessatisfy Bellman conditionsOptimal policydetermined from optimalvalues
How do Bellman optimality conditions connect action comparison, consistent values, and an optimal policy?
  1. Begin with the requirement that the values represent the largest expected returns achievable by any policy.
  2. Use the Bellman optimality equations as consistency conditions for those values.
  3. Solve the conditions in principle to obtain the optimal value functions.
  4. Use the resulting optimal values to determine an optimal policy.

Common Reasoning Errors

  • Treating a value as a permanent score attached to a state.

    A value function's meaning depends on the policy being followed.

    Fix: Always state which policy is being evaluated before interpreting a value.

  • Calling the value of one policy the optimal value.

    A policy value evaluates one specified policy, while an optimal value is the largest expected return achievable by any policy.

    Fix: Separate the question of policy performance from the question of best achievable performance.

  • Assuming that unique optimal values imply one unique optimal policy.

    For a given MDP, optimal values are unique, but multiple policies may share them.

    Fix: Treat the optimal value as fixed while allowing more than one policy to attain it.

  • Using Bellman optimality equations as though they only evaluate a chosen policy.

    They are consistency conditions for optimal value functions.

    Fix: Understand their role as conditions that can be solved for optimal values, followed by policy determination.

Check Your Understanding

MEDIUM

A learner says: “The value of state S is 8, so state S always has value 8.” Explain why this statement is incomplete. Then contrast the meaning of the value of S under one specified policy with the meaning of the optimal value of S.

Hints
  • Ask what behavior the ordinary value function assumes.
  • Ask whether the optimal value considers one policy or all policies.
  • Remember that optimal values can be shared by multiple optimal policies.

What do you think happens?

If two different policies achieve the same optimal values for a given MDP, must one of them be non-optimal?

  • Yes, because only one policy can be optimal.
  • No, multiple policies may attain the same optimal values.
  • It cannot be determined from the value functions.
Reveal answer

Answer: No, multiple policies may attain the same optimal values.

The optimal values for a given MDP are unique, but the policies that achieve them do not have to be unique.

Key Takeaways

  1. A policy's value functions assign expected returns to states and to state-action pairs under that policy.
  2. Ordinary value functions evaluate one specified policy.
  3. Optimal value functions select the largest expected return achievable by any policy.
  4. Optimal values can determine optimal policies, and several optimal policies may share the same optimal values.
  5. Bellman optimality equations provide consistency conditions that can be solved for optimal values before determining an optimal policy.

Key Takeaways

  • Value functions are policy-dependent expected-return evaluations.
  • State values begin from a state, while state-action values begin from a specified state-action pair.
  • Optimal value functions represent the largest expected return achievable by any policy.
  • An optimal policy can be determined from optimal values, and more than one policy may attain the same optimal values.
  • Bellman optimality equations state consistency conditions for finding optimal value functions.