Every-Visit MC
First-Visit MC keeps only the return after a state's first occurrence in each episode.
One Episode, Several Visits
Monte Carlo policy evaluation learns a state-value function from complete episodes generated by a policy π. It does not use a model of the environment. Instead, it observes what return followed a state and uses those observed returns to estimate the state's value. The central question for this article is what to do when the same state appears more than once in one episode.
Selecting Returns from an Episode
Suppose an episode encounters state s at several time steps. Each occurrence has a return that follows it. First-Visit MC selects only the return after the first occurrence of s in that episode. Every-Visit MC selects the returns after all occurrences of s in that episode. The selection happens separately for each state, so a later visit to s does not affect the First-Visit sample for s, even though it may still be selected by Every-Visit MC.
Repeated state in one episode
An episode visits state s three times. The returns following those visits are G₁, G₂, and G₃. Which returns enter the record for s?
Locate the visits: The episode contains three occurrences of s, in time order.
Apply First-Visit MC: Keep only G₁, the return following the first occurrence of s. This episode contributes at most one First-Visit sample for s.
Apply Every-Visit MC: Keep G₁, G₂, and G₃. Every occurrence contributes its following return.
For this episode, First-Visit MC records one return for s, while Every-Visit MC records three returns for s.
From Episode to Value Estimate
The evaluation process has two separate parts. First, the method selects returns from an episode according to its visit rule. Second, it averages the selected returns collected across episodes. For First-Visit MC, each episode contributes at most one return to the list for a particular state. Once that return has been appended, the value estimate V(s) is the average of the accumulated list.
The same separation clarifies the difference between First-Visit and Every-Visit MC. The averaging operation is not the defining difference: both methods use returns from episodes generated by π, and the stored returns are averaged. The difference is which returns are allowed into the stored list before averaging.
A Running Average of Samples
Imagine collecting First-Visit returns for state s across several complete episodes. Each new first visit supplies one more sample for the list associated with s. The current V(s) is the average of all those stored samples. A new episode can therefore change the estimate, but it does not replace the previous evidence.
Updating the stored evidence
Across three episodes, the first visit to state s produces returns G₁, G₂, and G₃.
Collect: Append the first-visit return from each episode to the stored list for s.
Retain: The list now contains the three selected returns. Earlier returns remain part of the evidence.
Estimate: Compute V(s) as the average of the stored returns.
The value estimate is the average of G₁, G₂, and G₃. The method estimates the state value from observed returns rather than from an environment model.
Why More Visits Help
Every additional First-Visit sample gives the average more information about the returns that follow state s under policy π. As more first visits are collected, the estimate moves toward vπ(s), the value associated with that policy. The source describes the error's standard deviation as decreasing as 1/√n, where n represents the number of collected samples. Thus, additional samples reduce variability, although the reduction becomes slower as the sample count grows.
First-Visit versus Every-Visit
Both methods learn from complete episodes generated by policy π, and both use observed returns rather than a model of the environment. Their difference is how repeated appearances of a state are counted. First-Visit MC selects one return per state per episode. Every-Visit MC uses every return following a visit to that state. Consequently, repeated visits create one First-Visit sample but potentially several Every-Visit samples within the same episode.
| Question | First-Visit MC | Every-Visit MC |
|---|---|---|
| What happens when a state appears repeatedly in one episode? | Keep only the return after its first occurrence. | Use the return after every occurrence. |
| How many samples can one episode provide for one state? | At most one. | One for each visit. |
| What happens after selection? | Average the stored returns. | Average the stored returns. |
The methods differ in return selection, not in the basic idea of averaging selected returns.
Mistakes with Repeated States
Treating First-Visit MC as if it recorded every occurrence of a state.
First-Visit MC keeps only the return after the state's first occurrence in that episode.
Fix:
Keep the first-occurrence return for First-Visit MC. Use all three returns only when applying Every-Visit MC.Thinking the distinction changes the averaging step.
The important boundary is return selection. Once selected returns are stored, the value estimate is produced by averaging the accumulated list.
Fix:
Ask which returns entered the list before asking how the list is averaged.Believing that one First-Visit sample means one sample total.
The first-visit rule applies within each episode. More episodes can provide more first-visit returns.
Fix:
Collect one qualifying return for s from each episode in which s occurs, then average the accumulated returns.Assuming more samples make every individual return equal to vπ(s).
The estimate approaches vπ(s) through the average of collected returns, while individual episode returns can differ.
Fix:
Track the running average and its decreasing variability rather than comparing each sample directly with vπ(s).
Check Your Understanding
An episode visits state s at time steps 2, 5, and 8. The returns following those visits are G₂, G₅, and G₈. State s also appears in two later episodes, once in each episode. Describe the complete set of returns that First-Visit MC records for s and the complete set that Every-Visit MC records for s.
Hints
- For First-Visit MC, inspect only the first occurrence of s within each episode.
- For Every-Visit MC, inspect every occurrence of s within each episode.
- Do not confuse the number of visits in one episode with the number of episodes.
What do you think happens?
If state s appears three times in one episode, how many First-Visit samples can that episode contribute for s?
Reveal answer
Answer: One
First-Visit MC keeps only the return following the first occurrence of s in that episode. Every-Visit MC would use all three following returns.
Key Takeaways
- First-Visit MC learns state values from complete episodes generated by policy π by averaging observed returns.
- When a state appears repeatedly in one episode, First-Visit MC records only the return after its first occurrence.
- Every-Visit MC records the return after every occurrence of the state.
- The evaluation sequence is episode generation, visit identification, return selection, return storage, and averaging.
- More first visits provide more returns for the average, moving the estimate toward vπ(s); the source describes the error's standard deviation as decreasing as 1/√n.
Key Takeaways
- First-Visit MC selects one return per state per episode, using the return after the state's first occurrence.
- Every-Visit MC uses every return following every visit to the state.
- After selection, both approaches use averages of stored returns to form V(s).
- As more first-visit samples are collected, the average becomes less variable and approaches vπ(s).