Concepts / Recursive Relationships of Value Functions

Recursive Relationships of Value Functions

The action-value function qπ(s, a) evaluates a specific action in a specific state under a policy.

  • Programming
Interactive lab

Try it: The Call Stack

How the call stack keeps track of function calls: each call pushes a frame with its own arguments and local variables, and each return pops it.

How it works

  1. Calling a function pushes a new frame on top of the stack; the caller pauses.
  2. The frame holds that call's parameters and local variables.
  3. When the function returns, its frame is popped and the return value goes back to the paused caller.
  4. A recursive function pushes one frame per call until the base case, then the frames unwind in reverse order.

Default run (17 steps): Program starts: main is about to call factorial(4). … The stack is empty again. Result: 24. Deepest point: 4 frames, 4 calls in total.

Educational simulation

Loading the simulation…

The Decision Hidden Inside a Value

A value estimate answers a question about future outcomes. The question becomes more specific when an action is included. Instead of asking only how valuable it is to be in state s, reinforcement learning can ask: if the agent is in state s and chooses action a, how valuable is that choice when policy π governs what happens afterward? That more specific quantity is the action-value function qπ(s, a).

The action-value function qπ(s, a) evaluates a particular action a taken in a particular state s, with policy π governing what happens afterward.

where action is takendecision being evaluatedgoverns what followsState sqπ(s, a)value of the chosen actionAction aPolicy π
How are a state, a selected action, and the policy connected when determining qπ(s, a)?

The action argument is the key distinction. qπ(s, a) does not evaluate the state alone; it evaluates a particular decision made in that state.

Separating Returns by Action

An action value can be estimated from what actually happens after the agent makes the decision. Each time the agent encounters state s and takes action a while following policy π, it observes the return that follows. The agent keeps the returns associated with that particular state-action pair and averages them. This produces an estimate of qπ(s, a).

Averaging Returns for One State-Action Pair

Suppose an agent encounters state s and chooses action a three times. The observed returns after those decisions are 6, 2, and 4. How can the agent estimate qπ(s, a)?

Collect matching experiences: Keep only the returns that followed action a when the agent was in state s: 6, 2, and 4.

Average the observed returns: Combine the returns associated with this same state-action pair and divide by the number of such observations.

Interpret the estimate: The resulting average is an experience-based estimate of the value of taking action a in state s while policy π governs what happens afterward.

The estimate is 4, obtained from the average of the three observed returns.

experience 1experience 2experience 3combinecombinecombineproduces(s, a)same state-action pairReturn 6Averageqπ(s, a)estimateReturn 2Return 4
How do multiple observed returns for the same state-action pair flow into an estimated action value?

State Values and Action Values

State-value estimation and action-value estimation organize experience differently. To estimate the state-value function vπ(s), the agent can average returns after encountering state s, regardless of which action was taken. To estimate qπ(s, a), the agent separates those experiences by action. If state s has several possible actions, each action receives its own average.

averageaveragevπ(s)returns after state sReturns after sactions are not separatedqπ(s, a)returns after s and aReturns after (s, a)each action has its owngroup
What information is averaged or conditioned on when estimating a state value compared with an action value?
EstimateExperiences grouped togetherInformation retained
vπ(s)Returns after encountering state sThe value of being in the state under policy π
qπ(s, a)Returns after taking action a in state sThe value of a particular decision in the state under policy π

The action argument determines whether experiences are separated by action.

Keeping action-specific averages preserves information that a state-only average would lose. Combining returns from different actions may describe the state, but it cannot separately describe the value of each action.

Tracing Value Through Successors

Value functions are not merely unrelated lists of estimates. They obey a recursive consistency relationship. The value assigned to a current state must be consistent with the rewards and the values associated with possible successor states reached after acting from that state. Evaluating the present therefore requires a connection to what can follow it.

What do you think happens?

An agent is evaluating a current state. If taking an action can lead to successor states with different values, should the current state's value be unrelated to those successor values?

  • Yes, because each state is evaluated independently
  • No, the current value must be consistent with what can follow it
  • Only when the state has one possible action
Reveal answer

Answer: No, the current value must be consistent with what can follow it.

The recursive relationship connects the value of the present to the rewards and values associated with possible successor states reached after acting.

takemay lead tomay lead tohas a valuehas a valueconstrains consistencyCurrent state svalue being evaluatedAction aSuccessor state s′one possible futureSuccessor valuesconsistency with whatfollowsSuccessor state s″another possible future
How does the value of the current state relate recursively to rewards and the values of successor states reached after taking actions?

The diagram does not say that every successor has the same value. It shows that the current evaluation must account for the possible futures reached after acting. This recursive consistency is a foundation used throughout reinforcement learning and dynamic programming.

Monte Carlo Estimates and Scaling Choices

The return-averaging approach belongs to the family of Monte Carlo methods. In this setting, Monte Carlo means estimating values by averaging many random samples of actual returns. The method waits for experience to produce returns and then uses those observed outcomes to improve estimates of states or state-action pairs.

For a problem with a manageable number of states or state-action pairs, maintaining a separate average for each one can preserve detailed information. When there are very many states, keeping a separate table of averages for every state or state-action pair may not be practical. An agent can instead represent vπ and qπ as parameterized functions with fewer parameters than there are states, then adjust those parameters so the functions better match observed returns.

RepresentationUseful whenMain ideaImportant consideration
Separate action averagesThe problem can support separate records for state-action pairsKeep the returns for each pair and average them independentlyPreserves action-specific information
Parameterized functionThere are very many states or state-action pairsUse fewer parameters to represent many values and adjust them using observed returnsAccuracy depends substantially on the chosen function approximator

Mistakes in Value Bookkeeping

  • Treating qπ(s, a) as a value of the state alone

    The action-value function evaluates a particular action in a particular state under a policy.

    Fix: Keep the state and action together when collecting returns for qπ(s, a).

  • Averaging returns from different actions into one action value

    Action-specific information is lost, so the result cannot separately represent the value of each action.

    Fix: Maintain a separate average for each action available in the state when estimating action values.

  • Assuming one observed return is the action value

    The practical Monte Carlo estimate is based on averaging many sampled actual returns.

    Fix: Use repeated experience with the same state-action pair to improve the estimate.

  • Treating value estimates as unrelated entries

    Value functions obey a recursive consistency relationship connecting present values with rewards and successor-state values.

    Fix: Check whether the current value is consistent with what can follow after acting.

Practice: Choose the Right Estimate

MEDIUM

An agent repeatedly encounters state s. Some experiences follow one action, and other experiences follow a different action. You need an estimate for the value of the first action specifically. Should you combine all returns after state s, or separate the returns by action? Explain what information would be lost by the other choice.

Hints
  • Look at whether the requested quantity includes an action argument.
  • State-value estimation can combine returns after encountering the state, while action-value estimation separates experiences by action.
MEDIUM

A problem has so many states that maintaining a separate average for every state-action pair is impractical. What representation choice described in this lesson can handle many values with fewer parameters, and what concern must you keep in mind?

Hints
  • Consider a representation that stands for many state or state-action values.
  • The quality of its estimates depends substantially on the chosen parameterized function approximator.

Key Takeaways

  • qπ(s, a) evaluates taking action a in state s while policy π governs what happens afterward.
  • A practical action-value estimate comes from averaging observed returns for the same state-action pair across repeated encounters.
  • State-value estimation groups returns by state, whereas action-value estimation separates returns by action within the state.
  • Value functions must be recursively consistent with rewards and the values of possible successor states.
  • Separate averages preserve detailed action information, while parameterized functions can represent many values when a separate table is impractical; their quality depends on the chosen approximator.