Concepts / Estimating Action Values with Monte Carlo Methods

Estimating Action Values with Monte Carlo Methods

A deterministic policy repeatedly selects the same action from a given state.

  • Programming

The Coverage Problem

Monte Carlo methods learn from complete observed returns. That sounds sufficient at first: let the agent follow its policy, observe what happens, and improve the estimates. However, a deterministic policy always selects the same action whenever it encounters a particular state. If several actions are available, repeatedly choosing only one of them can leave the other state-action pairs without any observed returns.

What do you think happens?

An agent reaches the same state three times, and its policy is deterministic. How many different actions will the policy select from that state?

  • Always the same action
  • A different action on every visit
  • An unpredictable number of actions
Reveal answer

Answer: Always the same action

A deterministic policy repeatedly selects the same action from a given state.

deterministic policyState sAction Aselected each time
When the agent reaches the same state multiple times, which action does the deterministic policy select each time?

Tracing State-Action Coverage

A state and an action from that state form a state-action pair. The policy can still reach the state while directing all experience toward only one available action. The selected pair receives observed returns. The other available action remains part of the choice problem, but no return from taking that action is available for its Monte Carlo estimate.

policy selectsnot selectedState savailable actionsAction AvisitedAction Bunvisited
Which actions are taken or skipped as the agent follows a fixed policy through a state?

Why Estimates Stall

Monte Carlo estimates improve by averaging observed returns. If an action is never selected, there are no returns from that action to average. Therefore, its estimate does not improve with the agent's experience, even though the action is still available from the state. This is why following only the action the policy currently favors is not enough to learn reliable values for every alternative.

producesaveraged intoreceivesState-action pairselected actionObserved returnsavailableMonte Carlo estimatecan improveState-action pairunselected actionNo observed returnsno improving estimate
How does experience flow into the estimate for a state-action pair, and what happens when an action is never selected?

When evaluating alternatives, ask whether each relevant action has observed returns behind its estimate. A policy's repeated preference for one action does not by itself establish the value of the alternatives.

From Returns to State Value

Monte Carlo state-value estimation takes a direct empirical route. Follow a fixed policy through complete episodes. Whenever the agent visits the state of interest, observe what happens afterward and record the resulting return. Each recorded return is one piece of evidence about the state's value under that policy. The state value is an expected return, while each recorded return is an observation used to estimate it.

observe afterwardrecordcombine across visitsVisit state swithin an episodeComplete outcomeafter the visitObserved returnone observationAverage returnstate-value estimate
How are rewards collected after visiting a state combined to estimate that state's value?

Evidence for a State's Value

An agent follows policy π through several complete episodes. After visiting state s, it records the return that follows the visit.

Visit: Each occurrence of state s is a possible source of evidence about the value of s.

Observe: The agent waits for the episode's outcome and records the resulting return.

Estimate: The recorded returns are averaged according to the selected visit rule.

The average of the selected observed returns is used as the estimate of the state's value under policy π.

Counting Visits in an Episode

A visit occurs whenever the state appears during an episode. A state may appear once, twice, or more in the same episode. Each occurrence is a visit, but first-visit and every-visit Monte Carlo use those occurrences differently. The first occurrence has a special role only because first-visit Monte Carlo deliberately keeps that occurrence and excludes later occurrences from the same episode.

next time stepnext time stepnext time stepcomplete episodeTime 0other stateTime 1state s: visitTime 2other stateTime 3state s: visitEpisode endreturns determined
At which time steps in an episode does the agent count as having visited a particular state?

Choosing the Averaging Rule

MethodReturns included when state s appears twicePurpose
First-visit Monte CarloThe return after the first occurrence onlyOne return per episode containing the state
Every-visit Monte CarloThe return after the first occurrence and the return after the later occurrenceEvery occurrence contributes evidence

Selecting Returns for One Episode

In one episode, state s appears twice. The return after the first occurrence is 5, and the return after the later occurrence is 3. Across another episode, a return of 1 follows an occurrence of s, and across another episode, a return of 7 follows an occurrence of s. Which returns are used by each method?

First-visit selection: Keep only the return after the first occurrence of s in each episode. The selected returns are 5 and 3.

Every-visit selection: Keep the return after every occurrence of s. The selected returns are 5, 1, 7, and 3.

Compare the sets: The difference is not the observed episode data. The difference is the rule that decides which occurrences contribute to the average.

First-visit MC uses 5 and 3. Every-visit MC uses 5, 1, 7, and 3.

includesincludesFirst-visit MC5 and 3one return per episodeEvery-visit MC5, 1, 7, and 3every occurrence
When a state appears multiple times in one episode, which returns are included in the first-visit and every-visit averages?

Mistakes in Visit Selection

  • Assuming that repeatedly following the favored action estimates every action value.

    No returns from taking Action B are available to improve its Monte Carlo estimate.

    Fix: Check whether each relevant state-action pair has observed returns before comparing their estimates.

  • Treating the state value as identical to one observed return.

    The state value is an expected return, while an individual return is only one observation.

    Fix: Use observed returns across visits and episodes to form an average estimate.

  • Including every occurrence when the method is first-visit Monte Carlo.

    First-visit MC includes only the return after the first occurrence in each episode.

    Fix: Discard later occurrences from that episode when applying the first-visit rule.

  • Using only the first occurrence when the method is every-visit Monte Carlo.

    Every-visit MC treats every occurrence as evidence.

    Fix: Include the return after each occurrence of the state.

Practice the Selection

EASY

State s appears twice in one episode. The return after the first occurrence is 5, and the return after the second occurrence is 3. Across two other episodes, the recorded returns after visits to s are 1 and 7. List the returns that belong in the first-visit estimate and the returns that belong in the every-visit estimate.

Hints
  • First-visit Monte Carlo keeps one return from each episode containing the state.
  • Every-visit Monte Carlo keeps the return after every occurrence of the state.
  • The source example uses the sets 5 and 3 for first-visit MC, and 5, 1, 7, and 3 for every-visit MC.
  1. To solve this kind of problem, identify every occurrence of the state, apply the selected visit rule, and only then determine which returns are averaged.

Key Takeaways

  • A deterministic policy selects the same action whenever it encounters a given state.
  • Following such a policy can leave other state-action pairs unvisited, so their Monte Carlo estimates do not improve from experience.
  • State-value estimation uses complete observed returns after visits to estimate an expected return.
  • First-visit Monte Carlo uses one return after the first occurrence of a state in each episode.
  • Every-visit Monte Carlo uses the return after every occurrence of the state.