State Values
An action value evaluates a state together with a particular action.
From Situations to Decisions
Suppose you know how valuable a situation is, but you do not know what each available action will do. That knowledge may not be enough to decide what to do next. State values answer the question, “How valuable is this state?” Reinforcement learning also needs the related question, “How valuable is each action from this state?”
An action value evaluates a state together with a particular action. Instead of assigning value only to a situation, an action value evaluates a state-action pair. Estimating these values gives information that can be used to suggest a policy because the values distinguish among the available actions from a state.
Choosing with a Model
State values alone can support a policy when a model is available. The model makes it possible to perform a one-step look-ahead decision: evaluate the possible next situations produced by each available action, then use their state values to select an action. In this setting, the value of a state helps guide the decision because the model supplies the missing connection between an action and the possible next state.
The model is essential in this use of state values. Without it, knowing the value of a possible successor state does not reveal which available action leads to that state. State values describe states, but they do not by themselves identify the quality of the actions that connect the current state to those states.
When the Model Is Missing
| Available information | What it supports |
|---|---|
| State values plus a model | A one-step look-ahead decision using possible next states |
| State values without a model | No direct identification of which action has the best outcome |
| Estimated values for each state-action pair | Information that can suggest which action to choose |
Learning from Complete Returns
Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome. Suppose an agent follows a fixed policy and repeatedly experiences episodes. To estimate the value of a particular state, record what happens after the agent visits that state and use the resulting returns as evidence about the state's value.
The state value is an expected return, while each recorded return is one observation used to estimate it. The estimation process averages the observed returns after visits to the state. As more returns are collected, their average should converge to the expected value.
For action values, the same Monte Carlo idea connects to the goal of estimating q ∗: collect returns associated with state-action pairs and use those observations to estimate how valuable the pairs are. Estimating action values is useful because it supplies action-specific information for suggesting a policy.
Counting Visits Correctly
A visit to a state is an occurrence of that state within an episode. A state can appear once, twice, or more in the same episode. Each occurrence is a visit, and a return can be recorded after each occurrence when using every-visit Monte Carlo estimation.
First-visit Monte Carlo includes one return per episode containing the state. It uses the return after the first occurrence of the state and ignores later occurrences in that same episode. Every-visit Monte Carlo includes a return after every occurrence of the state. If the episode returns to the state, both the first and later returns enter the every-visit average.
Worked Return Selection
Selecting Returns for One Repeated State
A state appears multiple times across the relevant episodes. The returns associated with the selected occurrences are 5 and 3 for the first-visit rule. The returns associated with all occurrences are 5, 1, 7, and 3. Which returns belong in each estimate?
Apply the first-visit rule: Select one return from each episode containing the state: the return after the first occurrence. The selected returns are 5 and 3.
Apply the every-visit rule: Select a return after every occurrence of the state. The selected returns are 5, 1, 7, and 3.
Match the average to the rule: The averaging step depends on the visit rule chosen before collecting the returns. Do not mix the two sets.
First-visit Monte Carlo uses 5 and 3. Every-visit Monte Carlo uses 5, 1, 7, and 3.
Choose the visit rule before averaging. When the rule is first-visit, keep only the first occurrence from each episode. When the rule is every-visit, keep the return after every occurrence. The correct estimate depends on selecting the appropriate returns.
Mistakes with Value Estimates
Treating a state value as if it directly identifies the best action.
Without a model, state values do not identify action quality by themselves.
Fix:
Use a model for one-step look-ahead, or estimate the values of the state-action pairs directly.Using only one observed return as the state's value.
The state value is an expected return, while each recorded return is only one observation used to estimate it.
Fix:
Collect observed returns and use their average as the estimate.Calling only the first occurrence a visit.
Each occurrence of the state within an episode is a visit.
Fix:
Include only the first occurrence for first-visit Monte Carlo, but include every occurrence for every-visit Monte Carlo.Mixing first-visit and every-visit return sets.
First-visit estimation uses one return per episode, whereas every-visit estimation uses a return after every occurrence.
Fix:
Select the returns according to the chosen visit rule before calculating the average.
Practice the Visit Rule
A state occurs twice in one episode and once in another. The returns after the occurrences in the first episode are 4 and 9, and the return after the occurrence in the second episode is 6. Which returns should be averaged for first-visit Monte Carlo? Which should be averaged for every-visit Monte Carlo?
Hints
- First-visit Monte Carlo keeps one return from each episode containing the state.
- Every-visit Monte Carlo keeps the return after every occurrence of the state.
The first-visit selection is 4 and 6. The every-visit selection is 4, 9, and 6. The difference is not in the observed episode data; it is in the rule used to decide which visits contribute evidence to the estimate.
Key Takeaways
- An action value evaluates a state together with a particular action.
- A model lets state values support a one-step look-ahead decision.
- Without a model, state values alone do not identify which action is best; action-value estimates provide action-specific information.
- Monte Carlo state-value estimation averages complete observed returns after visits to a state.
- First-visit Monte Carlo uses the first occurrence in each episode, while every-visit Monte Carlo uses every occurrence.
Key Takeaways
- State values describe the expected return of a state, while action values evaluate state-action pairs.
- A model connects candidate actions to possible next states, allowing state values to support one-step look-ahead.
- When no model is available, estimating each action's value gives more direct information for suggesting a policy.
- Monte Carlo estimation uses complete observed returns as evidence and averages them to estimate a state's value.
- First-visit and every-visit Monte Carlo differ in which returns are included when a state appears more than once in an episode.