Model Availability
An action value evaluates a state together with a particular action.
From Situation to Decision
Knowing that a situation is valuable does not automatically tell an agent what to do next. The missing information may be the effect of each available action. The central distinction in this topic is between evaluating a state by itself and evaluating a state together with a particular action.
A state value answers, in effect, how valuable a situation is. An action value evaluates a state-action pair: the current state together with one particular action.
One-Step Look-Ahead
When a model is available, state values can be used in a one-step look-ahead decision. The agent considers an available action, uses the model to determine what that action will do, and then uses the value of the possible next state when comparing actions. In this setting, state values can be enough to determine which action the policy should select because the model connects each candidate action to its possible next state.
Comparing two actions through their next states
An agent is in state S and can choose action A or action B. A model is available. The model indicates that A leads to state X and B leads to state Y. The known state value of X is higher than the known state value of Y.
Start with the current state: The agent begins in S and lists the actions available there.
Use the model: The model connects action A with X and action B with Y.
Use state values: The agent compares the value of X with the value of Y.
Select an action: Because X has the higher state value in this illustration, the comparison favors action A.
With a model available, state values can support a policy decision through one-step look-ahead.
When the Model Is Missing
When no model is available, knowing the value of the current state does not identify the quality of the actions available from that state. The agent may know that the current situation is valuable, but it does not know what each action will do. Without that action-to-outcome connection, the state value alone cannot determine which action should be selected.
Why one state value cannot rank actions
An agent knows that state S has a high value. From S, it can choose action A or action B, but no model is available to indicate what either action will do.
Record what is known: The value of S describes the situation as a whole.
Identify what is missing: There is no information connecting A or B to their possible outcomes.
Check whether the actions can be ranked: The single value for S does not say whether A or B is the better action.
The state value alone is insufficient for choosing between A and B when the model is unavailable.
Treating the value of the current state as the value of every available action.
A state value evaluates the situation, while an action value evaluates the situation together with a particular action.
Fix:
Ask whether the information distinguishes the available state-action pairs.Assuming that state values can always determine a policy.
The one-step look-ahead use of state values depends on having a model that connects actions with possible next states.
Fix:
Separate the model-available case from the model-unavailable case.Ignoring action identity when evaluating a state-action pair.
An action value requires both the state and a particular action.
Fix:
Represent the evaluation as a pair: current state plus selected action.
Action Values for Policy Suggestions
Estimating each available action's value makes the information useful for suggesting a policy. Instead of having only one value attached to the current state, the agent can compare the values associated with the different state-action pairs for that state. This comparison provides a direct basis for suggesting which action to take.
Comparing action values in one state
For the same state S, suppose an estimate is available for the pair consisting of S and action A, and another estimate is available for the pair consisting of S and action B. The estimate for the pair with A is higher.
Hold the state fixed: Both evaluations concern S, so the comparison is about different actions from the same situation.
Compare the action values: The estimate for the pair with A is compared with the estimate for the pair with B.
Suggest a policy choice: The higher estimated action value provides the stronger suggestion for the action to select.
Estimating action values turns action selection into a comparison among state-action pairs.
| Evaluation | What it includes | Use in this topic |
|---|---|---|
| State value | A state or situation | Can support action choice through one-step look-ahead when a model is available |
| Action value | A state together with a particular action | Allows action values for the same state to be compared directly |
Monte Carlo and q*
Monte Carlo methods connect with this topic through the goal of estimating q*. The relevant target is not only the value of a state, but the value associated with each state-action pair. Sampled episodes provide returns for those pairs, and repeated returns can improve the estimates used to compare actions. In this way, Monte Carlo estimation supports policy suggestions even when a model is not available.
What do you think happens?
An agent has no model, but it has estimates for the action values of two actions from the same state. Which information is more directly useful for suggesting the next action: one value for the state, or the two action-value estimates?
Reveal answer
Answer: The two action-value estimates
Action values evaluate the state together with particular actions, so comparing them directly distinguishes the available choices. This is the role connected with estimating q*.
Practice and Transfer
A learner says: “I know the current state has a high value, so I know which action to take.” Explain why this conclusion depends on model availability. Then describe what information would be more useful when no model is available.
Hints
- Ask whether the state value identifies what each action will do.
- Distinguish a state evaluation from a state-action evaluation.
- Explain how comparing estimated action values could suggest a policy.
When analyzing an action-selection problem, use this order: identify whether a model is available; determine whether the information evaluates states or state-action pairs; then ask whether the available values can be compared across the actions from the current state.
- Model availability changes how state values can be used. With a model, state values support one-step look-ahead because the model connects actions to possible next states. Without a model, the current state's value does not identify action quality by itself. Action values solve the comparison problem by evaluating each state together with a particular action. Estimating those values supports policy suggestions and connects Monte Carlo methods with the goal of estimating q*.
Key Takeaways
- An action value evaluates a state together with a particular action.
- A model allows state values to support a one-step look-ahead policy decision.
- Without a model, a state value alone does not identify which available action is better.
- Estimating action values makes it possible to compare actions from the same state.
- Monte Carlo methods connect with the goal of estimating q* for state-action pairs.