State-Value Estimation under a Policy
The methods differ only in how they select returns from episodes with repeated visits.
One Episode, Two Samples
Monte Carlo prediction estimates the value of a state under a policy by collecting episodes that follow that policy and using the returns associated with visits to the state. The interesting question is what to do when the same state appears more than once in one episode. First-visit and every-visit Monte Carlo use the same policy and aim at the same state value, vπ(s), but they select different numbers of returns from each episode.
The difference is a data-selection rule: first-visit records one return per episode for the state, while every-visit may record several returns from that episode.
Finding the First Occurrence
For first-visit Monte Carlo, inspect the episode from its beginning. If the target state appears once, record the return associated with that visit. If it appears several times, record only the return associated with the earliest occurrence and stop processing that state for this episode. Therefore, one episode contributes at most one return to the first-visit estimate for a given state.
Selecting One Return
In a generated episode, state s appears twice. The first occurrence has return 5, and the later occurrence has return 2. Which return enters the first-visit sample for s?
Scan from the beginning: The first occurrence of s is the visit with return 5.
Stop for this state: First-visit processing does not add the later occurrence from the same episode.
The episode contributes one first-visit return: 5.
Collecting Every Occurrence
Every-visit Monte Carlo follows the episode through all occurrences of the target state. Each occurrence contributes its associated return. If the state appears twice, the episode can contribute two returns; if it appears once, both methods contribute the same return from that episode.
Keeping Both Returns
Use the same generated episode: state s appears first with return 5 and later with return 2. Which returns enter the every-visit sample?
Process the first occurrence: Record return 5.
Continue through the episode: The later occurrence is also a visit to s, so record return 2 as well.
The episode contributes two every-visit returns: 5 and 2.
Side-by-Side Selection
Consider one generated episode in which s occurs at two points. The first occurrence is associated with return 5, and the second is associated with return 2. First-visit Monte Carlo places only 5 into its collection for s. Every-visit Monte Carlo places both 5 and 2 into its collection. The episode itself and the policy that generated it have not changed; only the selection rule has changed.
| Question | First-visit Monte Carlo | Every-visit Monte Carlo |
|---|---|---|
| How many returns can one episode contribute for a state? | One | Several |
| Which occurrence is used when the state repeats? | The earliest occurrence | Every occurrence |
| What happens when the state appears once? | The return is recorded | The same return is recorded |
| What grows as data is collected? | The number of first visits | The total number of visits |
Counting Without Distortion
Treating first-visit Monte Carlo as if it recorded every occurrence.
First-visit Monte Carlo records one return per episode for the state and stops after the earliest occurrence.
Fix:
Scan from the beginning and keep only the return associated with the first occurrence.Stopping after the first occurrence when using every-visit Monte Carlo.
Every-visit Monte Carlo continues through all occurrences and may record several returns from one episode.
Fix:
Process every occurrence of the state in the episode.Counting visits and returns inconsistently.
The estimate depends on the returns that actually enter the selected collection.
Fix:
After processing each episode, check that the visit count matches the number of returns recorded for the chosen method.Assuming the two methods use different policies.
Both methods use episodes obtained by following the same policy and aim to estimate the same state value.
Fix:
Separate the policy that generated the episode from the rule used to select returns.
When processing an episode, decide first whether the target state appears more than once. If it appears once, both methods record the same return. If it appears multiple times, use the earliest occurrence only for first-visit processing, and continue through all occurrences for every-visit processing.
More Visits, More Stability
A small number of returns can produce an estimate noticeably different from vπ(s). As more relevant visits are collected, the average becomes more stable. First-visit Monte Carlo approaches vπ(s) as the number of first visits increases. Every-visit Monte Carlo approaches vπ(s) as the total number of visits to s increases.
For first-visit Monte Carlo, the source characterizes each recorded return as an independent, identically distributed estimate of vπ(s) with finite variance. The law of large numbers implies that the average approaches the expected value as the number of first visits becomes very large. The standard deviation of the estimation error falls as 1 / √n, where n is the number of averaged returns; the source describes this as quadratic convergence. Every-visit Monte Carlo also converges to vπ(s), although its convergence argument is less straightforward because repeated returns within one episode are present.
What do you think happens?
Suppose you keep collecting episodes that follow the same policy. Which estimate eventually approaches vπ(s): first-visit, every-visit, both, or neither?
Reveal answer
Answer: Both
First-visit estimates converge as the number of first visits increases. Every-visit estimates converge as the total number of visits increases.
Practice the Selection Rule
A generated episode contains three occurrences of state s. Their associated returns, in episode order, are 4, 1, and 6. State s also appears in a second episode once, with return 3. List the returns that enter the first-visit collection and the returns that enter the every-visit collection. Then state how many returns each method has collected after these two episodes.
Hints
- For the first episode, first-visit keeps only the earliest occurrence.
- For the first episode, every-visit keeps all three occurrences.
- The second episode has only one occurrence, so both methods keep its return.
Practice Solution
Use the generated episodes with returns 4, 1, 6 for the repeated occurrences in the first episode, and return 3 for the single occurrence in the second episode.
First-visit collection: Keep 4 from the first episode and 3 from the second episode.
Every-visit collection: Keep 4, 1, and 6 from the first episode, then keep 3 from the second episode.
Count the samples: First-visit has 2 recorded returns. Every-visit has 4 recorded returns.
First-visit returns: 4 and 3. Every-visit returns: 4, 1, 6, and 3.
Essential Distinctions
- First-visit Monte Carlo records one return per episode for a given state, using the earliest occurrence.
- Every-visit Monte Carlo records the return from every occurrence of the state in the episode.
- If a state appears once, both methods record the same return from that episode.
- Both methods use episodes generated by the same policy and aim to estimate the same state value, vπ(s).
- With enough relevant data, both methods converge to vπ(s); first-visit counts first visits, while every-visit counts total visits.
Key Takeaways
- The central difference is how returns are selected when a state repeats within one episode.
- First-visit Monte Carlo keeps only the return from the earliest occurrence.
- Every-visit Monte Carlo keeps a return from each occurrence.
- Correct counting requires matching the number of recorded returns to the chosen selection rule.
- Both methods converge to vπ(s) as their relevant visit counts grow.