Concepts / Action-Value Functions

Action-Value Functions

The optimal state-value function v∗(s) gives the maximum expected return achievable from a state.

  • Programming

From Situation to Decision

In reinforcement learning, it is useful to evaluate both situations and decisions. A state tells us where the agent is, while an action tells us what the agent does from that state. Value functions describe the expected return associated with these possibilities.

The central distinction is whether the evaluation names only a state or also names a particular action. The optimal state-value function v∗(s) evaluates a state at the state level. The optimal action-value function q∗(s, a) evaluates a particular state-action pair.

v∗(s)stateq∗(s, a)state and action
How does evaluating a state differ from evaluating one specific action taken from that state?

Tracing an Initial Action

To interpret q∗(s, a), trace the situation in order. Begin in state s. Choose the particular action a. After that initial action, follow an optimal policy. The function evaluates the expected return for this complete description.

Reading q∗(s, a)

Interpret q∗(s, a) without assigning numerical returns.

Starting point: The description begins at state s. This identifies where the agent is before the decision.

Initial decision: The agent takes the specific action a. This action is fixed in the description rather than being left unspecified.

Continuation: After the initial action, the agent follows an optimal policy.

Evaluation: q∗(s, a) describes the expected return for starting at s, taking a first, and then following an optimal policy.

q∗(s, a) is an action-specific evaluation: it conditions on the starting state, the selected initial action, and optimal behavior afterward.

choosethen followevaluateState sstarting stateAction aselected first actionOptimal policyactions afterwardq∗(s, a)expected return
What happens after starting in state s, choosing action a, and then following the optimal policy?

State-Level Optimal Value

The optimal state-value function v∗(s) gives the maximum expected return achievable from state s.

The word maximum is important. v∗(s) asks how much return can be achieved from the state when the agent follows an optimal policy. It does not identify one particular first action in its notation. Instead, it gives a state-level view of the best expected return available from that state.

A useful way to read v∗(s) is: how much return can be achieved from this state when the agent follows an optimal policy? This differs from asking what happens after committing to one named action. The latter is the more specific question answered by q∗(s, a).

considerselect maximumState scurrent situationAvailable actionspossible decisionsMaximum expectedreturnv∗(s)
How does v∗(s) represent the maximum expected return available from a state across possible actions?

Comparing State and Action Values

FunctionWhat is fixed?What is evaluated?
v∗(s)Starting state sMaximum expected return achievable from the state
q∗(s, a)Starting state s and initial action aExpected return after taking a and then following an optimal policy
vπ(s)Starting state s and policy πExpected return from the state under policy π
qπ(s, a)Starting state s, initial action a, and policy π afterwardExpected return after taking a and thereafter following π

v∗(s) and q∗(s, a) are related but answer different questions. v∗ describes the best expected return at the level of a state. q∗ describes the expected return after the agent commits to a particular initial action from that state and then behaves optimally.

The policy-specific versions use the same distinction. A state-value function evaluates a state under a policy. The action-value function qπ(s, a) evaluates the expected return from starting at s, taking a, and thereafter following policy π.

policy-specific formpolicy-specific formv∗(s)state-level viewq∗(s, a)action-specific viewvπ(s)state under πqπ(s, a)action under π
How does evaluating a state differ from evaluating one specific action taken from that state?

Reading qπ(s, a) Precisely

The action-value function qπ(s, a) represents the expected return from starting in state s, taking action a, and thereafter following policy π.

A complete action-value description contains three fixed components. First, identify the starting state. Second, identify the action taken first. Third, identify the policy followed afterward. Omitting any of these can change what is being evaluated.

Three Fixed Components

Identify what remains fixed in the description qπ(s, a).

State: The starting state is s. It specifies where the evaluation begins.

Action: The selected action is a. It specifies the action taken before the policy controls what happens afterward.

Policy: The policy followed afterward is π. It specifies the policy used after the initial action.

qπ(s, a) evaluates one fixed starting state, one fixed initial action, and the subsequent behavior under policy π.

thenafterwardevaluateState sstarting stateAction aselected first actionPolicy πfollowed afterwardqπ(s, a)expected return
Which parts remain fixed in qπ(s, a): the starting state, the selected action, and the policy followed afterward?
beginfirstafterwardevaluates, a, πstate, action, policyStart at sTake aFollow πExpected returnqπ(s, a)
How is the expected return determined when the starting state and first action are fixed, but subsequent actions follow policy π?

Policy-Specific State Value

A state-value function evaluates a state under a policy. In the policy-specific notation vπ(s), the expected return is understood from state s while the agent follows policy π.

The policy matters because the value of a state is being described under a specified way of behaving. The same state-level question is different from the action-level question in qπ(s, a), where the first action is explicitly fixed before the policy is followed.

v∗(s)optimal policyvπ(s)policy π
How does the expected return from a state change when the agent follows a particular policy rather than choosing optimally?

Common Interpretation Errors

  • Treating v∗(s) as the value of one specific action.

    v∗(s) is indexed by a state, while q∗(s, a) is indexed by both a state and an action.

    Fix: Use q∗(s, a) when the initial action is part of the situation being evaluated.

  • Describing q∗(s, a) without mentioning the initial action.

    The selected action is essential to the action-value description.

    Fix: Say that the agent starts at s, takes a, and then follows an optimal policy.

  • Forgetting which policy is followed after the initial action.

    The policy followed afterward is one of the three essential parts of qπ(s, a).

    Fix: Name the subsequent policy explicitly.

  • Assuming that qπ(s, a) and vπ(s) describe exactly the same information.

    State value evaluates a state under a policy, whereas action value evaluates a particular initial action in that state followed by the policy.

    Fix: Check whether the specific first action is included.

Practice the Distinction

MEDIUM

For each description, identify whether it is a state-value or action-value description, and state what is fixed: (1) the expected return from state s under policy π, (2) the expected return after starting at s, taking action a, and then following π, (3) the maximum expected return achievable from state s.

Hints
  • Look for whether a particular initial action is named.
  • The notation v refers to a state-level evaluation, while q refers to an action-level evaluation.
  • The word optimal identifies the maximum expected return available from the state.

What do you think happens?

Which description includes more specific information: v∗(s) or q∗(s, a)?

  • v∗(s)
  • q∗(s, a)
  • They include exactly the same information
Reveal answer

Answer: q∗(s, a)

q∗(s, a) includes both the starting state and the particular initial action, while v∗(s) gives a state-level evaluation.

Key Takeaways

  1. v∗(s) gives the maximum expected return achievable from state s.
  2. q∗(s, a) evaluates a specific initial action from state s, followed by an optimal policy.
  3. v∗ provides a state-level view, while q∗ provides an action-specific view.
  4. vπ(s) evaluates a state under policy π.
  5. qπ(s, a) fixes the starting state, the selected initial action, and the policy followed afterward.

Key Takeaways

  • The optimal state-value function v∗(s) describes the maximum expected return available from a state.
  • The optimal action-value function q∗(s, a) describes the expected return after a specific initial action and optimal behavior afterward.
  • State value and action value differ because q includes a selected initial action while v does not.
  • The policy-specific function qπ(s, a) fixes the starting state, first action, and subsequent policy.
  • To interpret any value description correctly, track the state, action when present, and policy.