Returns and Episodes in Reinforcement Learning
The methods differ only in how they select returns from episodes with repeated visits.
One Episode, Two Data Rules
Monte Carlo prediction estimates the value of a state under a policy by collecting episodes that follow that policy and using the returns associated with visits to the state. The important question here is what to do when the same state appears more than once in one episode. First-visit Monte Carlo and every-visit Monte Carlo use the same policy and aim to estimate the same state value, vπ(s). They differ only in which visits contribute returns to the average.
First-visit Monte Carlo records one return per episode for a given state. Every-visit Monte Carlo may record several returns from one episode.
Tracing Visits Through an Episode
Consider an episode in which state S appears at two different time steps. For a generated example, suppose the return associated with the earlier occurrence is 8 and the return associated with the later occurrence is 4. These are two returns associated with visits to S in the same episode. The methods now make different selection decisions: one keeps only the earliest occurrence, while the other keeps both.
First-Visit Monte Carlo
First-visit Monte Carlo examines an episode from its beginning and finds the earliest occurrence of the state being estimated. It records the return associated with that first occurrence and does not record later occurrences of the same state from that episode. Therefore, one episode contributes at most one return for a given state.
- Choose the state whose value is being estimated.
- Inspect the episode in time order.
- Find the earliest visit to that state.
- Record the return associated with that earliest visit.
- Stop processing that state for this episode.
Selecting One Return
In a generated episode, state S appears at t=1 with return 8 and at t=3 with return 4. Which return does first-visit Monte Carlo record for S?
Find the earliest occurrence: The first occurrence of S is at t=1.
Use its associated return: The return associated with the first occurrence is 8.
Ignore later occurrences in this episode: The occurrence at t=3 does not add another return for first-visit processing.
The episode contributes one return for S: 8.
Every-Visit Monte Carlo
Every-visit Monte Carlo processes every occurrence of the state in the episode. If a state appears several times, the return associated with each occurrence enters the collection of returns used for the estimate. Consequently, one episode may contribute several returns for one state.
- Choose the state whose value is being estimated.
- Inspect every time step in the episode.
- Identify each visit to that state.
- Record the return associated with every identified visit.
- Continue until all occurrences in the episode have been processed.
Selecting Several Returns
In a generated episode, state S appears at t=1 with return 8 and at t=3 with return 4. Which returns does every-visit Monte Carlo record for S?
Process the first occurrence: The visit at t=1 contributes its associated return, 8.
Continue to the later occurrence: The visit at t=3 is also a visit to S, so it contributes its associated return, 4.
Collect both returns: The episode contributes both returns rather than stopping after the first occurrence.
The episode contributes two returns for S: 8 and 4.
Same Episode, Different Samples
| Question | First-visit Monte Carlo | Every-visit Monte Carlo |
|---|---|---|
| Which occurrences are processed? | The earliest occurrence of the state in the episode | Every occurrence of the state in the episode |
| How many returns can one episode contribute for one state? | One | Several |
| Generated episode with returns 8 and 4 | Records 8 | Records 8 and 4 |
| Convergence condition | The number of first visits becomes very large | The total number of visits becomes very large |
The distinction is a data-selection rule, not a policy change. Both methods use episodes obtained by following the same policy, and both aim to estimate vπ(s). Their difference is how many returns from each episode enter the average.
More Visits, More Stable Estimates
A small collection of returns can produce an estimate noticeably different from vπ(s). As more relevant visits are collected, the average generally becomes more stable. First-visit Monte Carlo approaches vπ(s) as the number of first visits increases. Every-visit Monte Carlo approaches vπ(s) as the total number of visits increases. Under the stated conditions, both methods converge to the same state value.
For first-visit Monte Carlo, the source describes the standard deviation of estimation error as falling at a rate of 1 / √n, where n is the number of averaged returns. Every-visit Monte Carlo also has quadratic convergence to vπ(s), although its theoretical justification is less direct because repeated returns within one episode are not treated in the same straightforward way.
Counting and Averaging Mistakes
Counting every occurrence when applying first-visit Monte Carlo
First-visit Monte Carlo records one return per episode for a given state and stops after the earliest occurrence.
Fix:
Keep only the return associated with the earliest occurrence, which is 8 in this generated example.Stopping after the first occurrence when applying every-visit Monte Carlo
Every-visit Monte Carlo continues through all occurrences of the state in the episode.
Fix:
Record the return associated with each occurrence, such as both 8 and 4.Treating the two methods as different policies
The distinction concerns which returns are selected from the episode, not the policy that generated it.
Fix:
Hold the policy and the episode-generation process conceptually constant, then apply the appropriate return-selection rule.Using the wrong number of returns in the average
The average must correspond to the returns actually collected under the selected method.
Fix:
For first-visit processing, count the recorded first visits. For every-visit processing, count all recorded visits.Expecting a small sample to equal vπ(s)
A small number of returns can produce an estimate noticeably different from the target value.
Fix:
Interpret the estimate as becoming more stable as more relevant visits are collected.
Practice the Selection Rule
A generated episode contains state S at three time steps. The associated returns are 7 for the first occurrence, 2 for the second occurrence, and 9 for the third occurrence. List the returns contributed by this episode under first-visit Monte Carlo and under every-visit Monte Carlo. Then state how many returns each method contributes.
Hints
- First identify the earliest occurrence of S.
- For first-visit processing, stop after that earliest occurrence.
- For every-visit processing, continue through all three occurrences.
What do you think happens?
For the practice episode, what returns should be recorded?
Reveal answer
Answer: First-visit: 7; every-visit: 7, 2, 9
First-visit Monte Carlo keeps only the return from the earliest occurrence. Every-visit Monte Carlo records the return from every occurrence of S.
The Essential Distinction
- First-visit Monte Carlo records the return from the first occurrence of a state in each episode.
- Every-visit Monte Carlo records the return from every occurrence of that state in each episode.
- If a state appears only once in an episode, both methods record the same return from that episode.
- The methods use the same policy-generated episodes and aim to estimate the same state value, vπ(s).
- As the relevant number of visits increases, both methods converge to vπ(s) under the stated conditions.
Key Takeaways
- First-visit Monte Carlo selects one return per episode for a given state by using the earliest occurrence.
- Every-visit Monte Carlo selects a return for every occurrence of the state in the episode.
- Repeated visits make the methods use different samples, even though both follow the same policy and target the same state value.
- The averaging denominator must match the number of returns actually recorded under the chosen method.
- Both methods become more stable and converge to vπ(s) as their relevant visit counts increase.