Learning State-Value Functions Under a Policy
Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome.
From Episodes to Evidence
Suppose an agent follows a fixed policy and repeatedly experiences episodes. You want to estimate how valuable a particular state is, but you do not begin with the value itself. Instead, you observe what happens after the agent reaches the state, record the resulting return, and use the collected returns as evidence.
The state value is an expected return. A recorded return is not the value itself; it is one observation used to estimate that value. Monte Carlo state-value estimation therefore learns from complete observed returns rather than from a single outcome. As more returns are collected after visits to the state, their average should converge to the expected value.
Pairing States with Returns
An episode supplies a sequence of states and a completed outcome. For each visit to a state, the return observed after that visit becomes evidence about the state's value. The important pairing is therefore state occurrence with the return that follows it, not state occurrence with a single immediate outcome.
A complete episode can contribute observed returns for the states visited during that episode. The estimate for a state is formed from the returns selected for that state under the chosen visit rule.
Recognizing a Visit
A visit occurs whenever the agent reaches the state during an episode. If state s appears once, the episode contains one visit to s. If state s appears twice or more, each occurrence is a visit to s. The occurrence's position in the episode matters because first-visit and every-visit methods select those occurrences differently.
Choosing the Visit Rule
| Method | Returns selected from an episode containing state s | Purpose of the selection |
|---|---|---|
| First-visit MC | One return: the return after the first occurrence of s | Limits each episode to its first visit to s |
| Every-visit MC | A return after every occurrence of s | Treats each occurrence of s as evidence |
First-visit Monte Carlo includes one return per episode containing the state. It deliberately uses only the return after the first occurrence of that state in the episode. Every-visit Monte Carlo includes a return after every occurrence of the state. If the episode returns to the state, both the return after the first occurrence and the return after the later occurrence enter the every-visit average.
Tracing a Repeated State
Selecting Returns for State s
An exercise provides the observed returns 5, 1, 7, and 3 for occurrences of a state across episodes. Which returns are used by first-visit Monte Carlo and which are used by every-visit Monte Carlo?
Apply the first-visit rule: First-visit Monte Carlo keeps one return per episode containing the state: the return after the first occurrence in that episode. In the exercise, the selected returns are 5 and 3.
Apply the every-visit rule: Every-visit Monte Carlo keeps the return after every occurrence of the state. In the exercise, all four observed returns enter the average: 5, 1, 7, and 3.
Compare the evidence sets: The two methods do not differ because the returns are calculated differently. They differ because they select different occurrences of the state before averaging.
First-visit MC uses 5 and 3. Every-visit MC uses 5, 1, 7, and 3.
The averaging step depends on the visit rule chosen before the returns are collected. For the source exercise, first-visit estimation averages the selected returns 5 and 3, whereas every-visit estimation averages 5, 1, 7, and 3.
Following the Episode
The complete process is empirical. A fixed policy generates episodes. Episodes contain visits to states. The outcomes after those visits provide observed returns. The chosen visit rule determines which of those returns are retained, and the retained returns are averaged to estimate the state value.
Mistakes in Return Selection
Treating the state value as identical to one observed return.
The state value is an expected return, while each recorded return is only an observation used to estimate it.
Fix:
Collect the relevant observed returns and use their average as the estimate.Counting an episode only once even when the state appears more than once.
Each occurrence of s is a visit, and every-visit MC uses a return after every occurrence.
Fix:
Record the occurrences separately, then apply either the first-visit or every-visit rule.Using every occurrence when first-visit MC was selected.
First-visit MC deliberately limits each episode to the return after the first occurrence of s.
Fix:
Keep only the first occurrence's return from that episode.Using only the first occurrence when every-visit MC was selected.
Every-visit MC treats each occurrence of s as evidence.
Fix:
Include the return after every occurrence of s.
Practice the Selection
A fixed policy generates episodes in which a state appears once in one episode and twice in another. For each method, describe which occurrences contribute returns to the estimate: first-visit Monte Carlo and every-visit Monte Carlo.
Hints
- For first-visit MC, select at most one return from each episode containing the state.
- For every-visit MC, select the return after each occurrence of the state.
- Do not decide from the numerical size of a return; decide from the occurrence selected by the method.
What do you think happens?
If one episode contains two visits to state s, how many returns from that episode can enter the estimate?
Reveal answer
Answer: First-visit MC can use one; every-visit MC can use two.
First-visit MC includes one return per episode containing the state, while every-visit MC includes a return after every occurrence of the state.
Working Rule
- Follow the fixed policy and observe complete episodes.
- Identify every occurrence of the state whose value is being estimated.
- Record the return observed after each occurrence.
- If using first-visit MC, retain only the return after the first occurrence in each episode.
- If using every-visit MC, retain the return after every occurrence.
- Average the retained returns to estimate the state's expected return.
The central decision is not which return looks most representative. It is which occurrences the selected Monte Carlo method allows into the estimate. First-visit MC uses one return per episode containing the state. Every-visit MC uses a return after every occurrence. In both cases, the retained observed returns provide evidence for the state's expected return.
Key Takeaways
- A state value is an expected return, while each recorded return is an observation used to estimate it.
- Every occurrence of a state within an episode counts as a visit.
- First-visit Monte Carlo includes one return per episode containing the state: the return after its first occurrence.
- Every-visit Monte Carlo includes the return after every occurrence of the state.
- The visit rule determines which observed returns are averaged.