Concepts / Understanding Monte Carlo Prediction

Understanding Monte Carlo Prediction

Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome.

  • Programming

From Visits to Value

Suppose an agent follows a fixed policy and repeatedly experiences episodes. You want to estimate how valuable a particular state is, but you do not begin with the value itself. Monte Carlo prediction takes a direct route: observe what happens after the agent reaches the state, record the resulting returns, and use their average as the estimate.

The state value is an expected return. Each recorded return is one observation used to estimate that expected value.

generatescontainsproducesaveraged with other returnsPolicy πEpisodeState s visitComplete returnValue estimate
How does a complete observed return flow from a state occurrence into an estimate of that state's value under the policy?

Following One State Through an Episode

An episode may contain a state once, twice, or more. Every occurrence of the particular state is a visit to that state. For each occurrence, look at what happens afterward and record the resulting return. That return is evidence about the value of the state under the policy being followed.

next stepnext stepnext stepTime 0state sTime 1another stateTime 2state sTime 3another state
What counts as a visit to a state, and where do those visits appear along the episode timeline?

Choosing Returns to Average

The visit rule determines which observations enter the estimate. First-visit Monte Carlo includes one return per episode containing the state: the return after the first occurrence. Every-visit Monte Carlo includes a return after every occurrence of the state. The two methods can therefore average different sets of returns from the same episodes.

First-visit MC5, 3Every-visit MC5, 1, 7, 3
How does the set of returns used for a state change when we average only the first occurrence versus every occurrence in each episode?
MethodReturn selected from an episodeTreatment of repeated visits
First-visit Monte CarloThe return after the first occurrence of the stateLater occurrences in the same episode are excluded
Every-visit Monte CarloThe return after every occurrence of the stateAll occurrences in the same episode contribute

The method must be chosen before deciding which returns to average.

A Return-Selection Walkthrough

Selecting returns for state s

A source exercise gives the return values 5, 1, 7, and 3 for occurrences of state s across episodes. The expected first-visit selection is 5 and 3.

First-visit selection: Use only the return after the first occurrence of s in each episode. The selected returns are 5 and 3.

Every-visit selection: Use the return after every occurrence of s. The selected returns are 5, 1, 7, and 3.

Compare the evidence: Both methods use observed returns, but they apply different rules to repeated occurrences within an episode.

First-visit MC averages 5 and 3. Every-visit MC averages 5, 1, 7, and 3.

first occurrencefirst occurrence in another episodeincludeincludeincludeincludeOccurrence 1return 5First-visit set5, 3Occurrence 2return 1Every-visit set5, 1, 7, 3Occurrence 3return 7Occurrence 4return 3
When the same state occurs at several time steps, which return follows each occurrence, and which of those returns should be averaged?

Accumulating Evidence Across Episodes

After applying the chosen visit rule to each episode, combine the selected returns by averaging them. The average is the Monte Carlo estimate for the state's value under the fixed policy. The key idea is empirical: each observed return is one piece of evidence, and as more returns are collected after visits to the state, their average should converge to the expected value.

contributecontributecontributeestimateEpisode 1selected returnAverageselected returns combinedState value estimateunder policy πEpisode 2selected returnMore episodesselected returns
How are the selected returns from multiple episodes combined to produce the Monte Carlo estimate for a state?

Decide whether you are using first-visit or every-visit Monte Carlo before collecting the returns to average. The same episode can provide different evidence sets under the two rules.

Mistakes with Repeated States

  • Treating a repeated state as only one visit regardless of the method.

    Every-visit Monte Carlo includes a return after every occurrence of the state.

    Fix: Ignore later occurrences only when using first-visit Monte Carlo.

  • Using every occurrence when the selected method is first-visit Monte Carlo.

    First-visit Monte Carlo includes one return per episode containing the state.

    Fix: Keep only the return after the first occurrence in each episode.

  • Confusing an observed return with the state value itself.

    The state value is an expected return, while a recorded return is one observation used to estimate it.

    Fix: Combine the selected observed returns by averaging them.

Practice the Selection Rule

MEDIUM

Use the source exercise's return values 5, 1, 7, and 3. Identify which returns belong in the first-visit estimate and which belong in the every-visit estimate. Then explain why the two selected sets differ.

Hints
  • First-visit Monte Carlo keeps one return per episode containing the state.
  • Every-visit Monte Carlo keeps the return after every occurrence of the state.
  • The source exercise identifies the first-visit selection as 5 and 3.

What do you think happens?

If an episode returns to state s, should the return after the later occurrence be included?

  • Always include it
  • Never include it
  • Include it for every-visit MC, but not for first-visit MC
Reveal answer

Answer: Include it for every-visit MC, but not for first-visit MC.

Every-visit Monte Carlo uses a return after every occurrence of the state, while first-visit Monte Carlo deliberately uses only the first occurrence in each episode.

Key Takeaways

  1. Monte Carlo state-value estimation uses complete observed returns to estimate the expected return for a state under a policy.
  2. Every occurrence of a state within an episode counts as a visit to that state.
  3. First-visit Monte Carlo uses one return per episode: the return after the first occurrence.
  4. Every-visit Monte Carlo uses the return after every occurrence of the state.
  5. The selected returns are averaged, and the average should converge toward the expected value as more returns are collected.

Key Takeaways

  • Monte Carlo prediction estimates a state's expected return from complete observed returns.
  • A state can be visited multiple times in one episode, and each occurrence is a visit.
  • First-visit MC selects only the first return for a state from each episode.
  • Every-visit MC selects a return after every occurrence of the state.
  • The chosen returns are averaged to produce the state-value estimate.