Concepts / Bellman Optimality Equation for q*

Bellman Optimality Equation for q*

v* assigns each state its maximum expected return.

  • Programming

From a State to Its Best Return

Suppose you are evaluating a state in which several actions are available. A policy would tell you which action choices to evaluate. The optimal value function takes a different viewpoint: it asks how much return can be obtained from that state when the best available action is used. The Bellman optimality equation expresses this idea as a consistency rule.

The optimal value function v* assigns each state its maximum expected return.

Tracing the State Value

consider available actionsproducesleads towardcombinecombinesets v*(s)State sv*(s)Best actionmaximum expected returnRimmediate rewardExpected returnR + γv*(s′)Next state s′γv*(s′)
How does the optimal value of a state combine the best action's immediate reward with the discounted expected value of the next state?
v*(s) = max over actions a of E[R + γv*(s') | s, a]

The expression being compared for an action is R + γv*(s'), evaluated in expectation given the state and action. R represents the reward in that expression, while γv*(s') represents the discounted optimal value associated with the next state. The maximization selects the largest expected return among the available actions.

Selecting the Maximum

considerconsiderlarger than 4not selecteddeterminesState savailable actionsAction 1expected return 7Action 1largest expected returnv*(s)7Action 2expected return 4
How do the possible actions at a state lead to different expected returns, and which action determines the state's optimal value?

Two actions from one state

A hypothetical state s has two possible actions. The expected return for the first action is 7, and the expected return for the second action is 4. What is the optimal value of s?

Evaluate each action: The first action has expected return 7, while the second action has expected return 4.

Apply maximization: The optimal value function retains the larger of the two action returns.

Identify the best action: The first action is the best action because its expected return is larger.

v*(s) follows the first action's expected return, so v*(s) is 7.

The numbers 7 and 4 illustrate the selection rule. The important point is that v*(s) is tied to the best action rather than to a policy action selected in advance.

Checking Bellman Consistency

test againstretain largestmust equalAssigned v*(s)value of the stateEvaluate actionsexpected R + γv*(s′)Maximum returnbest actionv*(s)agrees with maximum
How does an optimal state's value remain consistent with the rewards and optimal values that follow from choosing its best action?

The equation can be used as a consistency test. Start at a state, evaluate the expected return for every possible action, and retain the largest result. The resulting quantity must agree with the optimal value assigned to that state.

QuestionWhat is being selected?
Ordinary policy viewpointThe action choices specified by a policy
Optimal value viewpointThe action with the largest expected return

Mistakes in Maximization

  • Treating v*(s) as the return from any one action.

    The optimal value function assigns the state its maximum expected return.

    Fix: Evaluate the available action returns and retain the largest one.

  • Replacing the maximization with a preselected policy action.

    The Bellman optimality equation uses maximization so it does not need to name a particular policy.

    Fix: Compare the expected returns of the available actions and select the best result.

  • Ignoring the expected continuation value.

    The compared expression is R + γv*(s'), evaluated in expectation given a state and action.

    Fix: Include both the reward and the discounted optimal value associated with the next state when evaluating an action.

  • Failing the consistency check.

    The state value must agree with the expected return of the best action.

    Fix: Recompute the action returns and compare the assigned value with their maximum.

Practice and Summary

EASY

A state has three available actions with expected returns 5, 9, and 6. Identify the action return that determines v*(s), and explain why the other two values do not determine the optimal value.

Hints
  • The equation uses maximization over actions.
  • Compare all three expected returns.
  • The largest expected return determines the state's optimal value.
  1. v*(s) represents the maximum expected return available from state s. The Bellman optimality equation compares the expected expression R + γv*(s') for the available actions. Maximization selects the best action without requiring a particular policy to be named. The resulting maximum must agree with the optimal value assigned to the state.

Key Takeaways

  • v*(s) assigns a state its maximum expected return.
  • The Bellman optimality equation evaluates R + γv*(s') in expectation for each available action.
  • Maximization selects the best action rather than a preselected policy action.
  • The assigned optimal value must agree with the largest expected action return.