State-Value Estimation with Monte Carlo Methods
The target of estimation is qπ(s, a): the expected return after taking a in s and then following policy π.
From One Decision to an Estimate
Monte Carlo action-value estimation learns from episodes of experience. The target is qπ(s, a), the expected return after an agent takes action a in state s and then follows policy π. The central question is not merely whether the agent encountered state s. The question is what happened after the specific decision to take action a in that state.
For qπ(s, a), begin at the moment the agent is in s and takes a. The rest of that episode supplies an observed return that can be used as evidence about the target value.
Finding a Visit
A state-action pair is visited only when both parts occur together: the state is encountered and the specified action is taken there. Seeing state s by itself is not a visit to the pair (s, a). Likewise, taking action a in another state is not a visit to (s, a). After a visit, the remainder of the episode provides a return. If the same pair appears again later in that episode, that later occurrence is another possible visit with another following return.
Two Visits in One Episode
Consider this generated episode: at time 1 the agent is in s and takes a; later, at time 3, the agent is again in s and takes a.
Identify the first occurrence: The state and action match at time 1, so time 1 is a visit to (s, a).
Identify the repeated occurrence: The state and action match again at time 3, so time 3 is another visit to (s, a).
Attach returns: The remainder after time 1 supplies one return, and the remainder after time 3 supplies another return.
This episode contains two visits to (s, a), and therefore two possible following returns.
Following the Reward Sequence
Once a visit to (s, a) has been identified, look forward through the episode. The rewards obtained after that decision form the observed return associated with the visit. The initial state-action choice starts the observation, and the subsequent behavior under policy π determines what return is seen. That observed return is one piece of evidence about the expected return qπ(s, a).
Choosing Returns by Method
The difference between first-visit and every-visit Monte Carlo estimation is which visits within an episode contribute returns. First-visit Monte Carlo keeps only the return following the first visit to the state-action pair in that episode. Every-visit Monte Carlo keeps the returns following all visits to that pair in the episode.
| Situation in one episode | First-visit method | Every-visit method |
|---|---|---|
| The pair appears once | Use the return after that visit | Use the return after that visit |
| The pair appears twice | Use only the return after the first visit | Use the returns after both visits |
| The pair appears multiple times | Use only the return after the earliest visit | Use the return after every visit |
Return selection for one episode
Selecting Returns
In this generated episode, the pair (s, a) occurs three times. The return after the first occurrence is Return 1, the return after the second is Return 2, and the return after the third is Return 3.
Apply first-visit selection: The first-visit method includes Return 1 and excludes Return 2 and Return 3 for this episode.
Apply every-visit selection: The every-visit method includes Return 1, Return 2, and Return 3 for this episode.
Compare the evidence: Both methods use returns observed after visits, but they count repeated occurrences differently within the same episode.
First-visit contributes one return from this episode; every-visit contributes three.
Accumulating Evidence Across Episodes
Each selected return becomes evidence about qπ(s, a). Across episodes, the method collects the returns allowed by its visit rule and combines that growing set of observations into an estimate. First-visit uses at most one return for a given pair from each episode. Every-visit can use several returns from one episode when the pair recurs.
Convergence with More Visits
As visits to a state-action pair accumulate, both first-visit and every-visit Monte Carlo methods move toward the true expected action value qπ(s, a). The two methods differ in how they select evidence from an episode, but both use returns observed after visits to the pair. When the number of visits to each state-action pair approaches infinity, the source states that both methods converge quadratically to the true expected values.
| Method | Within-episode evidence | Long-run target |
|---|---|---|
| First-visit Monte Carlo | Return after the first visit in each episode | qπ(s, a) |
| Every-visit Monte Carlo | Returns after all visits in each episode | qπ(s, a) |
Practice: Classify the Evidence
A generated episode visits (s, a) at time 2 and again at time 5. The remainder after time 2 produces Return A, and the remainder after time 5 produces Return B. Which return or returns should be used by first-visit Monte Carlo? Which should be used by every-visit Monte Carlo?
Hints
- First identify the earliest occurrence of (s, a) in the episode.
- Then ask whether the method keeps only that occurrence or all occurrences.
What do you think happens?
Which returns are selected for the generated episode: first-visit uses Return A or Returns A and B, while every-visit uses Return A or Returns A and B?
Reveal answer
Answer: First-visit uses Return A; every-visit uses Returns A and B
Time 2 is the first visit in the episode, so first-visit keeps its following return. Every-visit keeps the returns following both occurrences.
Counting every occurrence of state s as a visit to (s, a).
A visit requires both the state and the specified action.
Fix:
Record a visit only when the agent is in s and takes a there.Using only the first return for every-visit estimation.
Every-visit Monte Carlo includes returns following all visits within the episode.
Fix:
Retain the return after each occurrence of the pair.Using all repeated returns for first-visit estimation.
First-visit Monte Carlo includes only the return following the first visit within the episode.
Fix:
Keep only the return associated with the earliest occurrence in that episode.Assuming the two methods estimate different action values.
Both methods target qπ(s, a), the expected return after taking a in s and then following policy π.
Fix:
Distinguish the methods by their return-selection rule, not by their target.
Summary
- qπ(s, a) is the expected return after taking action a in state s and then following policy π.
- A visit to (s, a) requires encountering state s and taking action a there.
- First-visit Monte Carlo uses the return after the first visit to the pair in each episode.
- Every-visit Monte Carlo uses the returns after all visits to the pair in each episode.
- As visits accumulate, both methods converge to the true expected action values; the source describes this convergence as quadratic.
Key Takeaways
- Action-value estimation learns qπ(s, a) from returns observed after taking a in s and following policy π.
- Only a state-action occurrence counts as a visit to the pair.
- First-visit keeps one return per pair per episode, while every-visit keeps returns after all occurrences.
- Repeated occurrences can produce distinct returns because each occurrence has a different episode remainder.
- With infinitely many visits, both methods converge to the true expected action values.