State-Value and Action-Value Functions
An optimal value function evaluates the best achievable outcome from a state.
The Question Behind a Value
A value function answers a question about expected return. The question changes depending on whether we are evaluating only a state or evaluating a state together with a particular first action. An optimal value function asks: starting in this state, what is the maximum value that can be obtained? A state-value function under a policy asks: starting in this state and following this policy, what value should we expect?
The word optimal means that available choices are considered and actions achieving the best possible result are followed.
Reading qπ(s, a)
The action-value function qπ(s, a) evaluates the expected return from a specific sequence of commitments. Begin in state s. Take action a as the first action. After that first action, follow policy π. The state, the first action, and the later policy are all essential to the description.
In the golf example, suppose the first stroke has already been fixed as a driver stroke. q∗(s, driver) evaluates the state under that commitment: the driver must be used first, while later strokes can be selected optimally from the available choices of driver or putter. The value therefore depends on more than the starting state; it also depends on the named first stroke.
State Value Versus Action Value
| Function | What is evaluated | What happens after the starting point |
|---|---|---|
| State value under policy π | State s | Follow policy π |
| Action value qπ(s, a) | State s and a particular first action a | Follow policy π after taking a |
| Optimal state value v∗(s) | The best achievable value from state s | Choose actions that achieve the best possible result |
| Optimal action value q∗(s, a) | State s after a particular first action a | Choose later actions optimally |
State value evaluates a state under a policy. It does not isolate one first action; the policy governs the choices that follow. Action value is more specific. It evaluates taking a particular action in the state and then following the policy afterward. Both descriptions involve a policy, but qπ(s, a) additionally names the initial action.
Bellman Optimality Checks
Bellman optimality equations characterize optimal values by relating the value of each state to the choices available from that state. They provide state-by-state relationships rather than treating the value function as an isolated list of numbers.
The recycling robot example has two states: high and low. Because there are two states, the Bellman optimality description contains two equations, one for the value of the high state and one for the value of the low state. Each equation must be consistent with the choices available in its corresponding state. Solving the system gives the optimal values for both states.
Think of the Bellman equations as mutual checks: each state's assigned value must agree with the choices available in that state, while those choices may depend on values of other states.
From Values to a Policy
Once the optimal value function has been solved, it can guide action selection. An optimal policy selects actions that achieve the maximum value from each state. The policy is therefore obtained by looking at the available actions and identifying which ones attain the optimal result.
In the gridworld example, solving the Bellman equation produces an optimal value for each state. The corresponding optimal policy identifies actions that achieve the maximum value from each state. If a gridworld cell contains multiple arrows, those arrows indicate that multiple actions are optimal there.
Following a Fixed Policy
A state-value function under a policy evaluates what happens when an agent starts in state s and continues following policy π. Unlike an optimal value function, this description is tied to the specified policy. It asks for the value of the state under that way of choosing actions, rather than the best value available across all choices.
Classifying a Value Description
Interpret the description: start in state s, take action a, and thereafter follow policy π.
Find the starting point: The description begins in state s.
Find the first commitment: The description names action a as the action taken first.
Find the later rule: After action a, the agent follows policy π.
Name the function: Because a specific first action is included, this is an action-value description: qπ(s, a).
qπ(s, a) evaluates the expected return from starting at s, taking a, and then following π.
Common Interpretation Errors
Treating qπ(s, a) as if it evaluated only the state.
The action-value function evaluates taking the particular action a first and then following π.
Fix:
Always identify the state, the first action, and the policy followed afterward.Confusing an optimal state value with a fixed first action.
The optimal state value concerns the best achievable value from the state, not a commitment to one named first action.
Fix:
Use q∗(s, a) when the first action is fixed; use v∗(s) for the best achievable value from the state.Assuming an optimal policy must contain exactly one action per state.
More than one action may achieve the maximum value.
Fix:
Allow every action that attains the maximum value to be part of the optimal policy.Treating a policy value as automatically optimal.
A state-value function evaluates a state under a specified policy, whereas an optimal value function evaluates the best achievable outcome.
Fix:
Check whether the description specifies a policy only or asks for optimal choices.
Practice: Identify the Evaluation
For each description, identify whether it is a state-value or action-value description, and name the fixed components. A. Start in state s and follow policy π. B. Start in state s, take action a first, and then follow policy π. C. From state s, choose actions that achieve the best possible result. D. The first golf stroke is fixed as a driver, while later strokes are selected optimally.
Hints
- Look for whether a specific first action is named.
- For an action-value description, list the starting state, first action, and later policy.
- The word optimal indicates that the best achievable result is being considered.
What do you think happens?
Which description is more specific: evaluating state s under policy π, or evaluating qπ(s, a)?
Reveal answer
Answer: Evaluating qπ(s, a)
Both descriptions involve a policy, but qπ(s, a) also fixes the initial action a.
Summary
- An optimal value function evaluates the best achievable outcome from a state.
- A state-value function evaluates a state under a policy.
- qπ(s, a) evaluates the expected return from starting at s, taking a first, and then following π.
- Bellman optimality equations relate each state's optimal value to the choices available from that state.
- An optimal policy selects actions that achieve maximum value, and multiple actions may be optimal.
Key Takeaways
- State value evaluates a state under a policy, while action value evaluates a particular first action in that state followed by a policy.
- The three fixed parts of qπ(s, a) are the starting state, the selected first action, and the policy followed afterward.
- An optimal state value represents the best achievable outcome from a state.
- Bellman optimality equations express state-by-state consistency among optimal values and available choices.
- An optimal policy selects maximizing actions, with ties allowing more than one optimal action.