Returns and Expected Discounted Rewards
Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome.
From State to Evidence
Suppose an agent follows a fixed policy and repeatedly experiences complete episodes. You want to estimate how valuable a particular state is, but the value is not given directly. Monte Carlo state-value estimation takes a direct route: whenever the agent reaches the state, observe what happens afterward, record the resulting return, and use observed returns as evidence about the state's value.
The state value is an expected return. An individual recorded return is not the value itself; it is one observation used to estimate that value. As more returns are collected after visits to the state, their average should converge to the expected value under the policy being followed.
Tracing One Visit
Begin at a particular visit to state s inside an episode. Look at the rewards that occur after that visit and use the resulting complete return as the observation associated with that visit. The return is therefore tied to a position in the episode: starting from an earlier visit gives the outcome of everything that follows from that point, while starting from a later visit gives the outcome from the later point.
A return is an observation obtained from what happens after a visit. Monte Carlo estimation uses complete observed returns, rather than trying to estimate the state value from a single outcome.
Repeated Visits in One Episode
A state can appear once, twice, or more in one episode. Each occurrence of state s is a visit to s. The occurrence's position matters because the return associated with that visit begins from that occurrence and reflects what follows it in the episode.
When an episode returns to s, the return after the first occurrence and the return after the later occurrence are separate observations for every-visit Monte Carlo. First-visit Monte Carlo deliberately keeps only the return after the first occurrence in that episode.
Choosing the Averaging Rule
Selecting returns for state s
The observed returns relevant to state s are 5 and 3 under the first-visit rule. Under the every-visit rule, the observed returns are 5, 1, 7, and 3. Which returns belong in each estimate?
First-visit selection: Use one return per episode containing s: the return after the first occurrence of s. The selected returns are 5 and 3.
Every-visit selection: Use a return after every occurrence of s. The selected returns are 5, 1, 7, and 3.
Apply the rule before averaging: The averaging step depends on the visit rule chosen. First decide which observations qualify, then average that selected collection.
First-visit MC averages 5 and 3. Every-visit MC averages 5, 1, 7, and 3.
Choose the visit rule before collecting or averaging observations. Otherwise, you may accidentally mix first-visit and every-visit data and produce an estimate that follows neither method.
Mistakes in Visit Selection
Treating the state value as if it were one observed return.
The value is an expected return, while an individual return is only one observation used for estimation.
Fix:
Collect observed returns and use their average according to the selected Monte Carlo visit rule.Counting a state only once because it is one distinct state.
Each occurrence of s is a visit to s. Every-visit MC includes a return after every occurrence.
Fix:
Inspect the episode positions and apply either the first-visit or every-visit rule.Using every occurrence for first-visit MC.
First-visit MC includes one return per episode containing the state.
Fix:
Keep only the return after the first occurrence in each episode.Using only the first occurrence for every-visit MC.
Every-visit MC treats each occurrence of s as evidence.
Fix:
Include a return after every occurrence of s.
If an episode contains s only once, first-visit MC and every-visit MC select the same return from that episode. The distinction appears when the episode contains s more than once.
Check the Selection
An episode contains state s twice. The return after the first occurrence is 5, and the return after the later occurrence is 1. Another episode contains s once, with return 3. Which returns should be averaged for first-visit MC? Which should be averaged for every-visit MC?
Hints
- First-visit MC keeps one return per episode containing s.
- Every-visit MC keeps a return after every occurrence of s.
What do you think happens?
Which return sets should be selected?
Reveal answer
Answer: First-visit uses 5 and 3. Every-visit uses 5, 1, and 3.
The first-visit rule selects one return from each episode, using the first occurrence of s. The every-visit rule selects both returns from the first episode and the one return from the second episode.
Key Takeaways
- A state value is an expected return; each recorded return is an observation used to estimate it.
- A visit is each occurrence of a state within an episode.
- First-visit Monte Carlo uses one return per episode containing the state, taken after the first occurrence.
- Every-visit Monte Carlo uses a return after every occurrence of the state.
- Select the correct returns first, then average them to form the estimate.
Key Takeaways
- Monte Carlo state-value estimation learns from complete observed returns.
- The state value is an expected return, not any single observed return.
- Every occurrence of a state is a visit, even when the state appears multiple times in one episode.
- First-visit MC uses only the first occurrence per episode, while every-visit MC uses all occurrences.
- The visit rule determines which returns belong in the average.