Action-Value Functions
The optimal state-value function v∗(s) gives the maximum expected return achievable from a state.
From Situation to Decision
In reinforcement learning, it is useful to evaluate both situations and decisions. A state tells us where the agent is, while an action tells us what the agent does from that state. Value functions describe the expected return associated with these possibilities.
The central distinction is whether the evaluation names only a state or also names a particular action. The optimal state-value function v∗(s) evaluates a state at the state level. The optimal action-value function q∗(s, a) evaluates a particular state-action pair.
Tracing an Initial Action
To interpret q∗(s, a), trace the situation in order. Begin in state s. Choose the particular action a. After that initial action, follow an optimal policy. The function evaluates the expected return for this complete description.
Reading q∗(s, a)
Interpret q∗(s, a) without assigning numerical returns.
Starting point: The description begins at state s. This identifies where the agent is before the decision.
Initial decision: The agent takes the specific action a. This action is fixed in the description rather than being left unspecified.
Continuation: After the initial action, the agent follows an optimal policy.
Evaluation: q∗(s, a) describes the expected return for starting at s, taking a first, and then following an optimal policy.
q∗(s, a) is an action-specific evaluation: it conditions on the starting state, the selected initial action, and optimal behavior afterward.
State-Level Optimal Value
The optimal state-value function v∗(s) gives the maximum expected return achievable from state s.
The word maximum is important. v∗(s) asks how much return can be achieved from the state when the agent follows an optimal policy. It does not identify one particular first action in its notation. Instead, it gives a state-level view of the best expected return available from that state.
A useful way to read v∗(s) is: how much return can be achieved from this state when the agent follows an optimal policy? This differs from asking what happens after committing to one named action. The latter is the more specific question answered by q∗(s, a).
Comparing State and Action Values
| Function | What is fixed? | What is evaluated? |
|---|---|---|
| v∗(s) | Starting state s | Maximum expected return achievable from the state |
| q∗(s, a) | Starting state s and initial action a | Expected return after taking a and then following an optimal policy |
| vπ(s) | Starting state s and policy π | Expected return from the state under policy π |
| qπ(s, a) | Starting state s, initial action a, and policy π afterward | Expected return after taking a and thereafter following π |
v∗(s) and q∗(s, a) are related but answer different questions. v∗ describes the best expected return at the level of a state. q∗ describes the expected return after the agent commits to a particular initial action from that state and then behaves optimally.
The policy-specific versions use the same distinction. A state-value function evaluates a state under a policy. The action-value function qπ(s, a) evaluates the expected return from starting at s, taking a, and thereafter following policy π.
Reading qπ(s, a) Precisely
The action-value function qπ(s, a) represents the expected return from starting in state s, taking action a, and thereafter following policy π.
A complete action-value description contains three fixed components. First, identify the starting state. Second, identify the action taken first. Third, identify the policy followed afterward. Omitting any of these can change what is being evaluated.
Three Fixed Components
Identify what remains fixed in the description qπ(s, a).
State: The starting state is s. It specifies where the evaluation begins.
Action: The selected action is a. It specifies the action taken before the policy controls what happens afterward.
Policy: The policy followed afterward is π. It specifies the policy used after the initial action.
qπ(s, a) evaluates one fixed starting state, one fixed initial action, and the subsequent behavior under policy π.
Policy-Specific State Value
A state-value function evaluates a state under a policy. In the policy-specific notation vπ(s), the expected return is understood from state s while the agent follows policy π.
The policy matters because the value of a state is being described under a specified way of behaving. The same state-level question is different from the action-level question in qπ(s, a), where the first action is explicitly fixed before the policy is followed.
Common Interpretation Errors
Treating v∗(s) as the value of one specific action.
v∗(s) is indexed by a state, while q∗(s, a) is indexed by both a state and an action.
Fix:
Use q∗(s, a) when the initial action is part of the situation being evaluated.Describing q∗(s, a) without mentioning the initial action.
The selected action is essential to the action-value description.
Fix:
Say that the agent starts at s, takes a, and then follows an optimal policy.Forgetting which policy is followed after the initial action.
The policy followed afterward is one of the three essential parts of qπ(s, a).
Fix:
Name the subsequent policy explicitly.Assuming that qπ(s, a) and vπ(s) describe exactly the same information.
State value evaluates a state under a policy, whereas action value evaluates a particular initial action in that state followed by the policy.
Fix:
Check whether the specific first action is included.
Practice the Distinction
For each description, identify whether it is a state-value or action-value description, and state what is fixed: (1) the expected return from state s under policy π, (2) the expected return after starting at s, taking action a, and then following π, (3) the maximum expected return achievable from state s.
Hints
- Look for whether a particular initial action is named.
- The notation v refers to a state-level evaluation, while q refers to an action-level evaluation.
- The word optimal identifies the maximum expected return available from the state.
What do you think happens?
Which description includes more specific information: v∗(s) or q∗(s, a)?
Reveal answer
Answer: q∗(s, a)
q∗(s, a) includes both the starting state and the particular initial action, while v∗(s) gives a state-level evaluation.
Key Takeaways
- v∗(s) gives the maximum expected return achievable from state s.
- q∗(s, a) evaluates a specific initial action from state s, followed by an optimal policy.
- v∗ provides a state-level view, while q∗ provides an action-specific view.
- vπ(s) evaluates a state under policy π.
- qπ(s, a) fixes the starting state, the selected initial action, and the policy followed afterward.
Key Takeaways
- The optimal state-value function v∗(s) describes the maximum expected return available from a state.
- The optimal action-value function q∗(s, a) describes the expected return after a specific initial action and optimal behavior afterward.
- State value and action value differ because q includes a selected initial action while v does not.
- The policy-specific function qπ(s, a) fixes the starting state, first action, and subsequent policy.
- To interpret any value description correctly, track the state, action when present, and policy.