Concepts / Optimal Policies in Reinforcement Learning

Optimal Policies in Reinforcement Learning

An optimal value function evaluates the best achievable outcome from a state.

  • Programming

The Best Possible Continuation

An optimal value function answers a focused question: starting in a particular state, what is the maximum value that can be obtained? It evaluates the best achievable outcome from that state after considering the available choices. The word optimal means that the later actions are selected to achieve the best possible result.

considerselect best resultCurrent stateavailable choicesPossible actionsconsideredOptimal valuemaximum achievable outcome
What does the optimal value function tell us about the best achievable expected return from a given state?

The optimal value belongs to the state-level question. It does not commit the agent to one named first action.

Reading State Values

Suppose an agent is in a particular state. Its optimal value describes the best result available from that point onward when the agent makes choices that achieve the best possible outcome. This value therefore summarizes the quality of the state under optimal decision-making. It is not merely the result of one predetermined action; it accounts for the choices available from the state.

The golf example makes this distinction concrete. An optimal value for a location answers how well the hole can be completed from that location when the available strokes are used in the best way. The value changes when the number of strokes available changes: a one-stroke contour covers only a small area near the hole, while two- and three-stroke contours reach progressively farther.

Interpreting a Golf State Value

Interpret an optimal value assigned to a golf state without fixing the first stroke.

Start at the state: Treat the ball's location as the state from which the agent must plan.

Consider available choices: The agent considers the available stroke choices rather than committing to one named first stroke.

Continue optimally: After each choice, later strokes can be selected to achieve the best possible result.

Interpret the value: The optimal value summarizes the best achievable outcome from that location under this optimal continuation.

An optimal state value evaluates the best result from the location as a whole; it is not the value of one specifically named first stroke.

more strokesmore strokesOne strokesmall area near holeTwo strokeslarger reachable regionThree strokesreaches the tee
How does the best achievable value change when the number of strokes available changes?

State Value versus Action Value

FunctionQuestion it answersFirst action
Optimal value functionWhat is the maximum value that can be obtained from this state?Not fixed to one named action
Optimal action-value functionWhat is the value of this state after a particular first action has been chosen?Fixed for the first action, with later choices made optimally

An optimal action-value function evaluates a state after a particular first action has been chosen. The first action is therefore part of the question. Later actions can still be selected optimally, but the initial commitment remains fixed.

In the golf example, q∗(s, driver) describes the value of state s under the commitment to use a driver first. The driver must be used for that first stroke, while later strokes may be selected optimally from the available driver or putter choices. By contrast, v∗(s) asks for the best achievable value from s without naming the first stroke in advance.

choose best continuationcontinue optimallyv∗(s)best from stateOptimal later actionsavailable after firstchoiceq∗(s, driver)driver fixed first
What is the difference between evaluating the best outcome from a state and evaluating the outcome of taking one specific action in that state?

Bellman Mutual Checks

Bellman optimality equations characterize optimal values by relating the value of each state to the choices available from that state. The equations express a state-by-state consistency requirement: an assigned optimal value must agree with the best choices and the resulting values of successor states.

The recycling robot example has two states: high and low. Because there are two states, the Bellman optimality description contains two equations, one for the high state and one for the low state. Each equation checks the value assigned to its state against the choices available there. Solving the equations gives the optimal values for both states.

inspectproducecompareassignCurrent statevalue to characterizeAvailable actionschoice setSuccessor statesresulting valuesMaximum valuebest action outcomeOptimal state valueconsistent assignment
How are optimal values propagated from successor states back to the current state through the maximum over actions?

Checking the Recycling Robot States

Explain how the Bellman optimality equations organize the two-state recycling robot example.

List the states: The state space contains high and low.

Write one relationship per state: The high state receives its own Bellman optimality equation, and the low state receives its own equation.

Account for available choices: Each equation relates its state's value to the choices available there and the resulting state values.

Solve the system: The two equations act as mutual checks. Solving them gives values that are consistent with optimal choices in both states.

Bellman optimality equations characterize the optimal values by requiring every state's value to agree with the best available choices.

From Values to Policy

Once the optimal value function has been obtained, it can be used to identify an optimal policy. For each state, the policy selects actions that achieve the maximum value. The policy is therefore the action-selection side of the value calculation: the values tell us how good the best attainable outcomes are, and the policy identifies actions that attain those outcomes.

evaluateevaluateevaluatecomparecomparecompareselect maximizersGridworld statecurrent cellNorthsuccessor valueMaximumbest action valueOptimal policyone or more actionsEastsuccessor valueSouthsuccessor value
Given the optimal values of possible successor states, how does an agent choose the action that defines the optimal policy?

In the gridworld example, solving the Bellman equation produces v∗, the optimal value function for the gridworld states. The corresponding optimal policy identifies actions that achieve the maximum value from each state. If several actions tie for the maximum, all of them can be optimal. This is why multiple arrows in one gridworld cell represent multiple optimal actions rather than an error.

comparecomparecompareselectAction AvalueMaximumhighest action valuePolicy actionmaximizing action oractionsAction BvalueAction Cvalue
How do the values of multiple available actions compare, and which action is selected as the maximizing choice?

Common Interpretation Errors

  • Treating the optimal value function as the value of one fixed first action.

    The optimal value function concerns the best achievable value from the state without a named first-action commitment.

    Fix: Use q∗(s, driver) when the first action is fixed as driver; use v∗(s) for the unrestricted best value from the state.

  • Assuming an optimal policy must select exactly one action in every state.

    More than one action may achieve the maximum value.

    Fix: Treat all maximizing actions as optimal when they tie.

  • Viewing Bellman equations as unrelated calculations for separate states.

    The equations form a system of mutual checks, with every state's value required to be consistent with the available choices.

    Fix: Consider one Bellman optimality equation for each state and solve them as a connected system.

Check Your Understanding

EASY

A state has several available actions, and two of them achieve the same maximum value. Explain what the optimal value function says about the state and what the optimal policy may contain.

Hints
  • Separate the state-level question from the action-selection question.
  • Remember that an optimal policy may contain more than one action in a state when actions tie for the maximum.
MEDIUM

In the golf example, explain why q∗(s, driver) is not the same question as v∗(s).

Hints
  • Identify whether the first stroke is fixed.
  • Later strokes may still be selected optimally in both interpretations.

Key Takeaways

  1. An optimal value function evaluates the best achievable outcome from a state.
  2. An optimal action-value function evaluates a state after a particular first action has been chosen.
  3. Bellman optimality equations relate each state's value to the available choices and characterize optimal values state by state.
  4. An optimal policy selects actions that achieve the maximum value.
  5. Several actions can be optimal in the same state when they tie for the maximum.

Key Takeaways

  • The optimal value function answers what maximum value can be obtained from a particular state.
  • The optimal action-value function answers what value results when a particular first action has already been chosen.
  • Bellman optimality equations provide mutually consistent, state-by-state relationships for optimal values.
  • An optimal policy is obtained by selecting actions that achieve the maximum value, including multiple tied actions when applicable.