Defining Optimal Policies
The optimal state-value function v∗(s) gives the maximum expected return achievable from a state.
Two Questions About Good Decisions
In reinforcement learning, it is useful to evaluate both situations and decisions. A state tells us where the agent is, while an action tells us what the agent does from that state. Optimal value functions describe the best expected return associated with these possibilities.
The central distinction is what the evaluation is conditioned on: v∗(s) is indexed by a state, whereas q∗(s, a) is indexed by both a state and an action.
Tracing the State-Level Evaluation
The optimal state-value function v∗(s) gives the maximum expected return achievable from state s. To interpret this definition, begin with the state and ask what return can be achieved when the agent follows an optimal policy from there. The value is not tied to one named first action; it summarizes the best expected outcome available from the state.
The word maximum is important. v∗(s) is not the expected return from an arbitrary policy or from an arbitrary action choice. It represents the greatest expected return achievable from that state when the agent follows an optimal policy.
Evaluating a Chosen First Action
The optimal action-value function q∗(s, a) evaluates a particular action in a particular state, followed by an optimal policy.
q∗(s, a) asks a more specific question than v∗(s): what return is expected if the agent first takes action a from state s and then follows an optimal policy? The initial action is fixed for this evaluation. The later behavior is evaluated as optimal.
Comparing Two Candidate Actions
An agent is in state s and can consider actions a and b. Explain what each function asks about this situation.
State-level question: v∗(s) asks how much return can be achieved from state s when the agent follows an optimal policy.
Action-specific question for a: q∗(s, a) asks for the expected return after the agent first takes action a in state s and then follows an optimal policy.
Action-specific question for b: q∗(s, b) asks the corresponding question for action b: take b first, then follow an optimal policy.
Relating the views: The action-specific evaluations allow the candidate first actions to be related to the best state-level evaluation. The exact numerical relationship depends on the reinforcement-learning setting.
v∗(s) provides the state-level maximum expected return, while q∗(s, a) and q∗(s, b) retain information about which initial action is being evaluated.
State Information Versus Action Information
| Function | Conditioned on | Question answered |
|---|---|---|
| v∗(s) | A state | How much return can be achieved from this state under an optimal policy? |
| q∗(s, a) | A state and a particular action | What return is expected after taking this action first and then following an optimal policy? |
This difference in indexing changes the information each function provides. v∗(s) tells us the maximum expected return associated with state s. q∗(s, a) tells us the expected return associated with the specific state-action pair, with the optimal-policy continuation included in the evaluation.
All optimal policies share the same optimal state-value function. Optimal action-value functions connect the action-specific evaluations in q∗ with the state-level evaluation in v∗.
Checking the First Decision
An agent is in state s. Describe, in words, the difference between asking for v∗(s) and asking for q∗(s, a). Then explain what happens after the initial action in the q∗ question.
Hints
- Check whether the question is indexed only by a state or by both a state and an action.
- For q∗, separate the fixed initial action from the later optimal-policy behavior.
What do you think happens?
Which function retains information about the particular initial action: v∗(s) or q∗(s, a)?
Reveal answer
Answer: q∗(s, a)
q∗ is indexed by both the state and the particular action. v∗ is indexed only by the state and gives a state-level view.
Common Interpretation Errors
Treating v∗(s) as the return from an arbitrary policy.
The optimal state-value function represents the maximum expected return achievable from the state under an optimal policy.
Fix:
Interpret v∗(s) as the best expected return available from that state.Treating q∗(s, a) as a state-only evaluation.
q∗ evaluates a particular action in a particular state.
Fix:
Keep both parts of the index in view: the state s and the initial action a.Assuming q∗ commits the agent to action a for all later decisions.
The function evaluates an initial action followed by an optimal policy.
Fix:
Describe the sequence as: start in state s, take action a, then follow an optimal policy.Assuming v∗ and q∗ answer identical questions.
They answer related but different questions. v∗ gives a state-level view, while q∗ gives an action-specific view.
Fix:
Ask whether the evaluation must distinguish among initial actions.
Essential Distinctions
- v∗(s) is the optimal state-value function and gives the maximum expected return achievable from state s.
- v∗(s) evaluates a state under an optimal policy without preserving a particular initial action in its index.
- q∗(s, a) is the optimal action-value function and evaluates a particular action in a particular state, followed by an optimal policy.
- The main difference is the information being conditioned on: v∗ uses a state, while q∗ uses a state-action pair.
- The two functions are related: action-specific q∗ evaluations connect to the optimal state-level evaluation v∗.
Key Takeaways
- v∗(s) gives the maximum expected return achievable from a state.
- q∗(s, a) evaluates a specified initial action in a specified state, followed by an optimal policy.
- v∗ provides a state-level view, while q∗ preserves action-specific information.
- To analyze q∗, separate the initial committed action from the optimal continuation.
- Optimal action-value evaluations connect the choice of an initial action with the optimal state-level evaluation.