Understanding Monte Carlo Prediction
Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome.
From Visits to Value
Suppose an agent follows a fixed policy and repeatedly experiences episodes. You want to estimate how valuable a particular state is, but you do not begin with the value itself. Monte Carlo prediction takes a direct route: observe what happens after the agent reaches the state, record the resulting returns, and use their average as the estimate.
The state value is an expected return. Each recorded return is one observation used to estimate that expected value.
Following One State Through an Episode
An episode may contain a state once, twice, or more. Every occurrence of the particular state is a visit to that state. For each occurrence, look at what happens afterward and record the resulting return. That return is evidence about the value of the state under the policy being followed.
Choosing Returns to Average
The visit rule determines which observations enter the estimate. First-visit Monte Carlo includes one return per episode containing the state: the return after the first occurrence. Every-visit Monte Carlo includes a return after every occurrence of the state. The two methods can therefore average different sets of returns from the same episodes.
| Method | Return selected from an episode | Treatment of repeated visits |
|---|---|---|
| First-visit Monte Carlo | The return after the first occurrence of the state | Later occurrences in the same episode are excluded |
| Every-visit Monte Carlo | The return after every occurrence of the state | All occurrences in the same episode contribute |
The method must be chosen before deciding which returns to average.
A Return-Selection Walkthrough
Selecting returns for state s
A source exercise gives the return values 5, 1, 7, and 3 for occurrences of state s across episodes. The expected first-visit selection is 5 and 3.
First-visit selection: Use only the return after the first occurrence of s in each episode. The selected returns are 5 and 3.
Every-visit selection: Use the return after every occurrence of s. The selected returns are 5, 1, 7, and 3.
Compare the evidence: Both methods use observed returns, but they apply different rules to repeated occurrences within an episode.
First-visit MC averages 5 and 3. Every-visit MC averages 5, 1, 7, and 3.
Accumulating Evidence Across Episodes
After applying the chosen visit rule to each episode, combine the selected returns by averaging them. The average is the Monte Carlo estimate for the state's value under the fixed policy. The key idea is empirical: each observed return is one piece of evidence, and as more returns are collected after visits to the state, their average should converge to the expected value.
Decide whether you are using first-visit or every-visit Monte Carlo before collecting the returns to average. The same episode can provide different evidence sets under the two rules.
Mistakes with Repeated States
Treating a repeated state as only one visit regardless of the method.
Every-visit Monte Carlo includes a return after every occurrence of the state.
Fix:
Ignore later occurrences only when using first-visit Monte Carlo.Using every occurrence when the selected method is first-visit Monte Carlo.
First-visit Monte Carlo includes one return per episode containing the state.
Fix:
Keep only the return after the first occurrence in each episode.Confusing an observed return with the state value itself.
The state value is an expected return, while a recorded return is one observation used to estimate it.
Fix:
Combine the selected observed returns by averaging them.
Practice the Selection Rule
Use the source exercise's return values 5, 1, 7, and 3. Identify which returns belong in the first-visit estimate and which belong in the every-visit estimate. Then explain why the two selected sets differ.
Hints
- First-visit Monte Carlo keeps one return per episode containing the state.
- Every-visit Monte Carlo keeps the return after every occurrence of the state.
- The source exercise identifies the first-visit selection as 5 and 3.
What do you think happens?
If an episode returns to state s, should the return after the later occurrence be included?
Reveal answer
Answer: Include it for every-visit MC, but not for first-visit MC.
Every-visit Monte Carlo uses a return after every occurrence of the state, while first-visit Monte Carlo deliberately uses only the first occurrence in each episode.
Key Takeaways
- Monte Carlo state-value estimation uses complete observed returns to estimate the expected return for a state under a policy.
- Every occurrence of a state within an episode counts as a visit to that state.
- First-visit Monte Carlo uses one return per episode: the return after the first occurrence.
- Every-visit Monte Carlo uses the return after every occurrence of the state.
- The selected returns are averaged, and the average should converge toward the expected value as more returns are collected.
Key Takeaways
- Monte Carlo prediction estimates a state's expected return from complete observed returns.
- A state can be visited multiple times in one episode, and each occurrence is a visit.
- First-visit MC selects only the first return for a state from each episode.
- Every-visit MC selects a return after every occurrence of the state.
- The chosen returns are averaged to produce the state-value estimate.