Estimating Action Values with Monte Carlo Methods
A deterministic policy repeatedly selects the same action from a given state.
The Coverage Problem
Monte Carlo methods learn from complete observed returns. That sounds sufficient at first: let the agent follow its policy, observe what happens, and improve the estimates. However, a deterministic policy always selects the same action whenever it encounters a particular state. If several actions are available, repeatedly choosing only one of them can leave the other state-action pairs without any observed returns.
What do you think happens?
An agent reaches the same state three times, and its policy is deterministic. How many different actions will the policy select from that state?
Reveal answer
Answer: Always the same action
A deterministic policy repeatedly selects the same action from a given state.
Tracing State-Action Coverage
A state and an action from that state form a state-action pair. The policy can still reach the state while directing all experience toward only one available action. The selected pair receives observed returns. The other available action remains part of the choice problem, but no return from taking that action is available for its Monte Carlo estimate.
Why Estimates Stall
Monte Carlo estimates improve by averaging observed returns. If an action is never selected, there are no returns from that action to average. Therefore, its estimate does not improve with the agent's experience, even though the action is still available from the state. This is why following only the action the policy currently favors is not enough to learn reliable values for every alternative.
When evaluating alternatives, ask whether each relevant action has observed returns behind its estimate. A policy's repeated preference for one action does not by itself establish the value of the alternatives.
From Returns to State Value
Monte Carlo state-value estimation takes a direct empirical route. Follow a fixed policy through complete episodes. Whenever the agent visits the state of interest, observe what happens afterward and record the resulting return. Each recorded return is one piece of evidence about the state's value under that policy. The state value is an expected return, while each recorded return is an observation used to estimate it.
Evidence for a State's Value
An agent follows policy π through several complete episodes. After visiting state s, it records the return that follows the visit.
Visit: Each occurrence of state s is a possible source of evidence about the value of s.
Observe: The agent waits for the episode's outcome and records the resulting return.
Estimate: The recorded returns are averaged according to the selected visit rule.
The average of the selected observed returns is used as the estimate of the state's value under policy π.
Counting Visits in an Episode
A visit occurs whenever the state appears during an episode. A state may appear once, twice, or more in the same episode. Each occurrence is a visit, but first-visit and every-visit Monte Carlo use those occurrences differently. The first occurrence has a special role only because first-visit Monte Carlo deliberately keeps that occurrence and excludes later occurrences from the same episode.
Choosing the Averaging Rule
| Method | Returns included when state s appears twice | Purpose |
|---|---|---|
| First-visit Monte Carlo | The return after the first occurrence only | One return per episode containing the state |
| Every-visit Monte Carlo | The return after the first occurrence and the return after the later occurrence | Every occurrence contributes evidence |
Selecting Returns for One Episode
In one episode, state s appears twice. The return after the first occurrence is 5, and the return after the later occurrence is 3. Across another episode, a return of 1 follows an occurrence of s, and across another episode, a return of 7 follows an occurrence of s. Which returns are used by each method?
First-visit selection: Keep only the return after the first occurrence of s in each episode. The selected returns are 5 and 3.
Every-visit selection: Keep the return after every occurrence of s. The selected returns are 5, 1, 7, and 3.
Compare the sets: The difference is not the observed episode data. The difference is the rule that decides which occurrences contribute to the average.
First-visit MC uses 5 and 3. Every-visit MC uses 5, 1, 7, and 3.
Mistakes in Visit Selection
Assuming that repeatedly following the favored action estimates every action value.
No returns from taking Action B are available to improve its Monte Carlo estimate.
Fix:
Check whether each relevant state-action pair has observed returns before comparing their estimates.Treating the state value as identical to one observed return.
The state value is an expected return, while an individual return is only one observation.
Fix:
Use observed returns across visits and episodes to form an average estimate.Including every occurrence when the method is first-visit Monte Carlo.
First-visit MC includes only the return after the first occurrence in each episode.
Fix:
Discard later occurrences from that episode when applying the first-visit rule.Using only the first occurrence when the method is every-visit Monte Carlo.
Every-visit MC treats every occurrence as evidence.
Fix:
Include the return after each occurrence of the state.
Practice the Selection
State s appears twice in one episode. The return after the first occurrence is 5, and the return after the second occurrence is 3. Across two other episodes, the recorded returns after visits to s are 1 and 7. List the returns that belong in the first-visit estimate and the returns that belong in the every-visit estimate.
Hints
- First-visit Monte Carlo keeps one return from each episode containing the state.
- Every-visit Monte Carlo keeps the return after every occurrence of the state.
- The source example uses the sets 5 and 3 for first-visit MC, and 5, 1, 7, and 3 for every-visit MC.
- To solve this kind of problem, identify every occurrence of the state, apply the selected visit rule, and only then determine which returns are averaged.
Key Takeaways
- A deterministic policy selects the same action whenever it encounters a given state.
- Following such a policy can leave other state-action pairs unvisited, so their Monte Carlo estimates do not improve from experience.
- State-value estimation uses complete observed returns after visits to estimate an expected return.
- First-visit Monte Carlo uses one return after the first occurrence of a state in each episode.
- Every-visit Monte Carlo uses the return after every occurrence of the state.