Recursive Relationships of Value Functions
The action-value function qπ(s, a) evaluates a specific action in a specific state under a policy.
Try it: The Call Stack
How the call stack keeps track of function calls: each call pushes a frame with its own arguments and local variables, and each return pops it.
How it works
- Calling a function pushes a new frame on top of the stack; the caller pauses.
- The frame holds that call's parameters and local variables.
- When the function returns, its frame is popped and the return value goes back to the paused caller.
- A recursive function pushes one frame per call until the base case, then the frames unwind in reverse order.
Default run (17 steps): Program starts: main is about to call factorial(4). … The stack is empty again. Result: 24. Deepest point: 4 frames, 4 calls in total.
Loading the simulation…
The Decision Hidden Inside a Value
A value estimate answers a question about future outcomes. The question becomes more specific when an action is included. Instead of asking only how valuable it is to be in state s, reinforcement learning can ask: if the agent is in state s and chooses action a, how valuable is that choice when policy π governs what happens afterward? That more specific quantity is the action-value function qπ(s, a).
The action-value function qπ(s, a) evaluates a particular action a taken in a particular state s, with policy π governing what happens afterward.
The action argument is the key distinction. qπ(s, a) does not evaluate the state alone; it evaluates a particular decision made in that state.
Separating Returns by Action
An action value can be estimated from what actually happens after the agent makes the decision. Each time the agent encounters state s and takes action a while following policy π, it observes the return that follows. The agent keeps the returns associated with that particular state-action pair and averages them. This produces an estimate of qπ(s, a).
Averaging Returns for One State-Action Pair
Suppose an agent encounters state s and chooses action a three times. The observed returns after those decisions are 6, 2, and 4. How can the agent estimate qπ(s, a)?
Collect matching experiences: Keep only the returns that followed action a when the agent was in state s: 6, 2, and 4.
Average the observed returns: Combine the returns associated with this same state-action pair and divide by the number of such observations.
Interpret the estimate: The resulting average is an experience-based estimate of the value of taking action a in state s while policy π governs what happens afterward.
The estimate is 4, obtained from the average of the three observed returns.
State Values and Action Values
State-value estimation and action-value estimation organize experience differently. To estimate the state-value function vπ(s), the agent can average returns after encountering state s, regardless of which action was taken. To estimate qπ(s, a), the agent separates those experiences by action. If state s has several possible actions, each action receives its own average.
| Estimate | Experiences grouped together | Information retained |
|---|---|---|
| vπ(s) | Returns after encountering state s | The value of being in the state under policy π |
| qπ(s, a) | Returns after taking action a in state s | The value of a particular decision in the state under policy π |
The action argument determines whether experiences are separated by action.
Keeping action-specific averages preserves information that a state-only average would lose. Combining returns from different actions may describe the state, but it cannot separately describe the value of each action.
Tracing Value Through Successors
Value functions are not merely unrelated lists of estimates. They obey a recursive consistency relationship. The value assigned to a current state must be consistent with the rewards and the values associated with possible successor states reached after acting from that state. Evaluating the present therefore requires a connection to what can follow it.
What do you think happens?
An agent is evaluating a current state. If taking an action can lead to successor states with different values, should the current state's value be unrelated to those successor values?
Reveal answer
Answer: No, the current value must be consistent with what can follow it.
The recursive relationship connects the value of the present to the rewards and values associated with possible successor states reached after acting.
The diagram does not say that every successor has the same value. It shows that the current evaluation must account for the possible futures reached after acting. This recursive consistency is a foundation used throughout reinforcement learning and dynamic programming.
Monte Carlo Estimates and Scaling Choices
The return-averaging approach belongs to the family of Monte Carlo methods. In this setting, Monte Carlo means estimating values by averaging many random samples of actual returns. The method waits for experience to produce returns and then uses those observed outcomes to improve estimates of states or state-action pairs.
For a problem with a manageable number of states or state-action pairs, maintaining a separate average for each one can preserve detailed information. When there are very many states, keeping a separate table of averages for every state or state-action pair may not be practical. An agent can instead represent vπ and qπ as parameterized functions with fewer parameters than there are states, then adjust those parameters so the functions better match observed returns.
| Representation | Useful when | Main idea | Important consideration |
|---|---|---|---|
| Separate action averages | The problem can support separate records for state-action pairs | Keep the returns for each pair and average them independently | Preserves action-specific information |
| Parameterized function | There are very many states or state-action pairs | Use fewer parameters to represent many values and adjust them using observed returns | Accuracy depends substantially on the chosen function approximator |
Mistakes in Value Bookkeeping
Treating qπ(s, a) as a value of the state alone
The action-value function evaluates a particular action in a particular state under a policy.
Fix:
Keep the state and action together when collecting returns for qπ(s, a).Averaging returns from different actions into one action value
Action-specific information is lost, so the result cannot separately represent the value of each action.
Fix:
Maintain a separate average for each action available in the state when estimating action values.Assuming one observed return is the action value
The practical Monte Carlo estimate is based on averaging many sampled actual returns.
Fix:
Use repeated experience with the same state-action pair to improve the estimate.Treating value estimates as unrelated entries
Value functions obey a recursive consistency relationship connecting present values with rewards and successor-state values.
Fix:
Check whether the current value is consistent with what can follow after acting.
Practice: Choose the Right Estimate
An agent repeatedly encounters state s. Some experiences follow one action, and other experiences follow a different action. You need an estimate for the value of the first action specifically. Should you combine all returns after state s, or separate the returns by action? Explain what information would be lost by the other choice.
Hints
- Look at whether the requested quantity includes an action argument.
- State-value estimation can combine returns after encountering the state, while action-value estimation separates experiences by action.
A problem has so many states that maintaining a separate average for every state-action pair is impractical. What representation choice described in this lesson can handle many values with fewer parameters, and what concern must you keep in mind?
Hints
- Consider a representation that stands for many state or state-action values.
- The quality of its estimates depends substantially on the chosen parameterized function approximator.
Key Takeaways
- qπ(s, a) evaluates taking action a in state s while policy π governs what happens afterward.
- A practical action-value estimate comes from averaging observed returns for the same state-action pair across repeated encounters.
- State-value estimation groups returns by state, whereas action-value estimation separates returns by action within the state.
- Value functions must be recursively consistent with rewards and the values of possible successor states.
- Separate averages preserve detailed action information, while parameterized functions can represent many values when a separate table is impractical; their quality depends on the chosen approximator.