Value Function Estimation
The action-value function qπ(s, a) evaluates a specific action in a specific state under a policy.
Why the Action Matters
Reinforcement learning often asks how valuable it is for an agent to be in a particular state. That question is answered by the state-value function, written vπ(s). But sometimes the state alone is not specific enough. The agent may have several possible actions, and each choice can lead to a different return. The action-value function qπ(s, a) addresses this more specific question: if the agent is in state s and chooses action a, how valuable is that choice when policy π governs what happens afterward?
The extra action argument is the central distinction: vπ(s) evaluates a state, while qπ(s, a) evaluates one particular action taken in that state under a policy.
Reading qπ(s, a)
To interpret qπ(s, a), keep the three ingredients separate. State s identifies where the agent is. Action a identifies the decision being evaluated. Policy π describes how behavior is governed after that decision. The resulting action value evaluates that particular state-action choice rather than the state in isolation.
One State, Two Action Choices
An agent is in state s and can choose actions a1 or a2. What should be kept separate when estimating the value of these choices?
Identify the state: Both decisions begin in the same state s.
Identify the action: The first choice is a1 and the second choice is a2. These are different state-action pairs even though the state is the same.
Track the policy: For each choice, the later behavior is governed by policy π.
Name the estimates: The two action-specific estimates are qπ(s, a1) and qπ(s, a2).
The action-value function preserves which action was selected in state s; a state-only value would not preserve that distinction.
From Returns to Estimates
An action value can be estimated from experience. Each time the agent encounters state s, takes action a, and observes the return that follows, it records that return with the same state-action pair. After repeated encounters, the agent averages the recorded returns. This average becomes an estimate of qπ(s, a).
Three Observed Returns
Suppose the same state-action pair (s, a) produces observed returns of 8, 4, and 6 on three different experiences. What estimate results from averaging them?
Collect the returns: Keep 8, 4, and 6 together because all three followed the same state-action pair (s, a).
Combine the observations: Add the observed returns to obtain 18.
Divide by the number of observations: There are three observations, so divide 18 by 3.
The estimated action value is 6 for this generated example.
State Values and Action Values
| Feature | State-value estimation | Action-value estimation |
|---|---|---|
| Quantity | vπ(s) | qπ(s, a) |
| What is evaluated | Being in a state under a policy | Taking a particular action in a state under a policy |
| How returns are grouped | By state, regardless of which action was taken | By state-action pair, keeping actions separate |
| Information preserved | The state identity | The state and the selected action |
For a state-value estimate, returns after encountering state s can be averaged without separating them by action. For an action-value estimate, returns must be separated by action. If state s has several possible actions, each action receives its own average. Combining all actions into one state-only average loses information about which decision produced each return.
Recursive Value Consistency
Value estimates are not merely unrelated numbers stored for unrelated situations. They obey recursive relationships. The value assigned to a state must be consistent with what can follow after acting from that state: the immediate reward and the values of the possible successor states. Evaluating the present therefore requires a connection to later states.
The diagram shows the direction of the consistency relationship without requiring a separate calculation formula. Starting from state s, an action produces an immediate reward and may lead to one of several successor states. The value of the current state must reflect both what is received immediately and the values associated with what can follow.
Choosing a Representation
Separate averages are useful when the agent can maintain an estimate for each state or state-action pair. However, problems with very many states may make a separate table impractical. In that situation, vπ and qπ can be represented by parameterized functions with fewer parameters than there are states. The parameters are adjusted so the functions better match observed returns.
Mistakes in Value Estimation
Treating qπ(s, a) as if it evaluated only the state.
The action argument is what distinguishes the value of one decision from the value of the state alone.
Fix:
Keep separate estimates for the relevant state-action pairs.Combining returns from different actions into one action-value average.
Action-value estimation groups observations by both state and action.
Fix:
Associate each observed return with the exact state-action pair that produced it.Assuming one observed return is enough.
Monte Carlo estimation uses averages of many sampled actual returns.
Fix:
Continue collecting repeated experiences and update the average.Assuming a large problem always fits a separate table.
Maintaining separate averages for every pair may not be practical at large scale.
Fix:
Consider a parameterized value function, while recognizing that its quality depends on the chosen approximator.
Check Your Understanding
An agent encounters state s three times. On the first encounter it takes action a1 and observes return 5. On the second it takes action a2 and observes return 9. On the third it takes action a1 and observes return 7. Which observations belong in the estimate for qπ(s, a1), and why?
Hints
- Group an observation by both the state and the selected action.
- Do not include the return produced after action a2 in the action-a1 estimate.
- The two returns associated with a1 can be averaged.
What do you think happens?
If the agent later encounters the same state s but selects a different action, should that return automatically be added to the existing qπ(s, a1) average?
Reveal answer
Answer: No, because the action is different.
Action-value estimation keeps returns separate by state-action pair. A return after another action belongs to that other action's estimate, not automatically to qπ(s, a1).
Key Takeaways
- qπ(s, a) evaluates a particular action in a particular state when policy π governs what happens afterward.
- Repeated observed returns for the same state-action pair can be averaged to estimate its action value.
- State-value estimation groups returns by state, while action-value estimation keeps returns separate by action.
- Value estimates must be recursively consistent with immediate rewards and the values of possible successor states.
- Separate averages can be practical for manageable problems; parameterized functions can help when there are very many states, but their quality depends on the chosen approximator.
Key Takeaways
- The action-value function qπ(s, a) evaluates a selected action in a state under a policy.
- A practical action-value estimate comes from averaging repeated returns observed after the same state-action pair.
- Keeping action-specific averages preserves information that state-only averaging loses.
- Current state values must be consistent with immediate rewards and possible successor-state values.
- Large state spaces may require parameterized functions instead of a separate average for every state-action pair.