Concepts / Policies and Actions

Policies and Actions

Value functions evaluate expected return under a particular policy.

  • Programming

From Behaviour to Evaluation

Suppose an agent is in a state and must decide what to do next. A policy tells the agent how to behave, but it does not by itself tell us how promising the current situation is. Value functions provide that evaluation: they describe the expected return associated with a state or with a state-action pair when the agent follows a particular policy.

A policy answers the question, “What should the agent do?” A value function answers, “How much expected return is associated with this state or state-action pair under that policy?”

evaluatesevaluatesPolicybehaviour ruleStateexpected returnState-action pairexpected return
What does each value function contain, and how does a state value differ from the value of taking a specific action in that state?

Two Kinds of Policy Value

A policy's value functions evaluate expected return under that particular policy. The state-value function evaluates a state: it describes the expected return associated with being in that state while following the policy. The action-value function evaluates a state-action pair: it describes the expected return associated with taking a particular action in that state while following the policy.

Evaluating a State and an Action

An agent is in state S. Its policy specifies how it behaves. What do the two value-function viewpoints evaluate?

State viewpoint: The state-value viewpoint evaluates state S under the particular policy. It asks how much expected return is associated with being in S when the policy governs behaviour.

State-action viewpoint: The action-value viewpoint evaluates a pair such as state S together with a particular action. It asks how much expected return is associated with that choice when the agent follows the policy.

Policy dependence: Both evaluations are tied to the particular policy being considered. They describe performance under that policy rather than the best performance available across every policy.

The state value evaluates the situation as a whole under a policy, while the action value evaluates a specified action in that situation under a policy.

Ordinary and Optimal Evaluation

Ordinary value functions answer: “How well does this policy perform?” They evaluate expected return while one particular policy is being followed. Optimal value functions ask a different question: “What is the best performance any policy could achieve?” They keep the largest expected return available across policies.

Value typeWhat it evaluatesPolicies considered
Policy value functionExpected return associated with a state or state-action pairOne particular policy
Optimal value functionLargest expected return availableAll available policies
evaluateskeeps largestOne policyexpected returnPolicy valueevaluationAll policiesavailable choicesOptimal valuelargest expected return
How do expected returns under one particular policy differ from the best possible expected returns across all policies?

Constructing an Optimal Policy

An optimal value function does more than report the best achievable expected return. It can also be used to construct an optimal policy through greedy choices. In each state, the policy selects an action associated with the highest available value. A policy that makes these choices is greedy with respect to the optimal values.

considerconsidercompare valuecompare valuechoose greedilyStateaction choicesAction AvalueLargest valuegreedy choiceOptimal policyselected actionAction Bvalue
Given the values of different actions in a state, how does selecting the highest-valued action produce an optimal policy?

Using Values to Choose

In one state, suppose the available actions have expected-return values of 8 and 5. Which action would a policy choose if it is greedy with respect to the optimal values?

Compare: The policy compares the expected-return values associated with the available actions.

Select: The action with value 8 has the larger value, so a greedy policy selects it.

Interpret: The selection is based on the optimal-value viewpoint: choose an action associated with the largest expected return available in that state.

The greedy policy selects the action with value 8.

Bellman Optimality Updates

Bellman optimality equations provide the mechanism for finding optimal value functions. Their role is to constrain value estimates so that the values are mutually consistent with optimal behavior. The update considers possible next states and rewards, then uses the best expected return available from the current situation to refine the optimal value estimate.

The important mental model is a two-stage route. First, the equations constrain the values until they are consistent with optimal behavior. Second, the resulting optimal values make policy selection relatively easy: choose a policy that is greedy with respect to those values.

considerlead toproduceevaluateevaluateupdateCurrent statevalue estimatePossible actionsavailable choicesNext statespossible outcomesBest expected returnlargest available valueOptimal valueupdated estimateRewardspossible returns
How does the Bellman optimality equation use possible next states and rewards to update an optimal value estimate?

When Optimal Policies Tie

An optimal value function can be unique even when the optimal policy is not. This happens when several different actions, or several different policies, produce the same largest expected return. Each policy can then be optimal because each achieves the optimal value, even though the policies make different choices.

choosechoosesupportssupportsStatedecision pointAction Amaximum valuePolicy Achooses AAction Bmaximum valuePolicy Bchooses B
How can two or more different actions or policies produce the same maximum value in a state?

Two Policies, One Optimal Value

Suppose two different actions available in the same state both produce the same largest expected return. What follows?

Identify the tie: Neither action has a larger expected return than the other. Both attain the largest value available in that state.

Construct choices: One policy can choose the first action, while another policy can choose the second action.

Compare policies: Although the policies differ in their choices, both achieve the same optimal value.

Multiple optimal policies can exist even though the optimal value function is unique.

Common Reasoning Errors

  • Treating a policy as if it already evaluates how promising a state is.

    A policy tells the agent how to behave. Value functions provide the evaluation of a state or state-action pair under that policy.

    Fix: Separate the behaviour rule from the value function that evaluates expected return.

  • Confusing a policy value with an optimal value.

    A policy's value functions evaluate one particular policy, while optimal value functions keep the largest expected return available across policies.

    Fix: Ask whether the evaluation concerns one specified policy or the best policy available.

  • Assuming that one unique optimal value requires one unique optimal policy.

    Several policies may be optimal even though the optimal value functions are unique.

    Fix: Check whether multiple choices attain the same maximum value.

  • Using the Bellman optimality equations only as a policy-selection rule.

    The equations help find optimal value functions by constraining values to be consistent with optimal behaviour.

    Fix: Think in two stages: find mutually consistent optimal values, then choose a policy greedily with respect to them.

Check Your Understanding

MEDIUM

An agent has a policy that determines how it behaves. Explain the difference between evaluating a state under that policy and evaluating a specific state-action pair under that policy. Then explain how the question changes when moving from ordinary value functions to optimal value functions.

Hints
  • State values evaluate a state, while action values evaluate a state together with a specified action.
  • Ordinary value functions concern one particular policy.
  • Optimal value functions keep the largest expected return available across policies.
EASY

Two different policies make different choices in a state, but both achieve the same largest expected return. Are both policies allowed to be optimal? Explain why.

Hints
  • Focus on the value achieved, not only on whether the choices match.
  • Several policies can be optimal when they attain the same maximum value.

The Decision Pipeline

  1. A policy specifies behaviour, while value functions evaluate expected return under that policy.
  2. State-value functions evaluate states; action-value functions evaluate state-action pairs.
  3. Ordinary value functions evaluate one particular policy, whereas optimal value functions keep the largest expected return available across policies.
  4. Optimal values can be used to construct an optimal policy by making greedy choices.
  5. Bellman optimality equations help find values that are mutually consistent with optimal behaviour.
  6. Multiple optimal policies can exist when different choices produce the same maximum value, even though the optimal value functions are unique.

Key Takeaways

  • Value functions evaluate expected return for states and state-action pairs under a particular policy.
  • Optimal value functions represent the largest expected return available across policies.
  • A policy that is greedy with respect to optimal values is an optimal policy.
  • Bellman optimality equations help constrain value estimates so they are consistent with optimal behaviour.
  • Several different policies may be optimal when they achieve the same maximum value.