Concepts / State-Value Estimation with Monte Carlo Methods

State-Value Estimation with Monte Carlo Methods

The target of estimation is qπ(s, a): the expected return after taking a in s and then following policy π.

  • Programming

From One Decision to an Estimate

Monte Carlo action-value estimation learns from episodes of experience. The target is qπ(s, a), the expected return after an agent takes action a in state s and then follows policy π. The central question is not merely whether the agent encountered state s. The question is what happened after the specific decision to take action a in that state.

For qπ(s, a), begin at the moment the agent is in s and takes a. The rest of that episode supplies an observed return that can be used as evidence about the target value.

takethen followepisode remainderevidence forState sAction aPolicy πObserved returnqπ(s, a)Expected return
How does the initial state-action decision connect to the return used to estimate qπ(s, a)?

Finding a Visit

A state-action pair is visited only when both parts occur together: the state is encountered and the specified action is taken there. Seeing state s by itself is not a visit to the pair (s, a). Likewise, taking action a in another state is not a visit to (s, a). After a visit, the remainder of the episode provides a return. If the same pair appears again later in that episode, that later occurrence is another possible visit with another following return.

next stepnext stepnext stepnext stept = 0state x, action at = 1state s, action at = 2state y, action bt = 3state s, action at = 4episode ends
At which time steps does an episode visit the specific pair (s, a), and how can repeated occurrences be identified?

Two Visits in One Episode

Consider this generated episode: at time 1 the agent is in s and takes a; later, at time 3, the agent is again in s and takes a.

Identify the first occurrence: The state and action match at time 1, so time 1 is a visit to (s, a).

Identify the repeated occurrence: The state and action match again at time 3, so time 3 is another visit to (s, a).

Attach returns: The remainder after time 1 supplies one return, and the remainder after time 3 supplies another return.

This episode contains two visits to (s, a), and therefore two possible following returns.

Following the Reward Sequence

Once a visit to (s, a) has been identified, look forward through the episode. The rewards obtained after that decision form the observed return associated with the visit. The initial state-action choice starts the observation, and the subsequent behavior under policy π determines what return is seen. That observed return is one piece of evidence about the expected return qπ(s, a).

thenthenuntilsupplies(s, a)visit timeRewardafter the visitRewardlater in episodeEpisode endReturnremainder after visit
After taking action a in state s, how does the remainder of the episode become the return used for estimation?

Choosing Returns by Method

The difference between first-visit and every-visit Monte Carlo estimation is which visits within an episode contribute returns. First-visit Monte Carlo keeps only the return following the first visit to the state-action pair in that episode. Every-visit Monte Carlo keeps the returns following all visits to that pair in the episode.

selectsselectsselectsFirst-visitkeep return after firstoccurrenceReturn 1includedEvery-visitkeep return after everyoccurrenceReturn 1includedReturn 2included
When the same state-action pair appears multiple times in one episode, which returns belong to each estimate?
Situation in one episodeFirst-visit methodEvery-visit method
The pair appears onceUse the return after that visitUse the return after that visit
The pair appears twiceUse only the return after the first visitUse the returns after both visits
The pair appears multiple timesUse only the return after the earliest visitUse the return after every visit

Return selection for one episode

Selecting Returns

In this generated episode, the pair (s, a) occurs three times. The return after the first occurrence is Return 1, the return after the second is Return 2, and the return after the third is Return 3.

Apply first-visit selection: The first-visit method includes Return 1 and excludes Return 2 and Return 3 for this episode.

Apply every-visit selection: The every-visit method includes Return 1, Return 2, and Return 3 for this episode.

Compare the evidence: Both methods use returns observed after visits, but they count repeated occurrences differently within the same episode.

First-visit contributes one return from this episode; every-visit contributes three.

Accumulating Evidence Across Episodes

Each selected return becomes evidence about qπ(s, a). Across episodes, the method collects the returns allowed by its visit rule and combines that growing set of observations into an estimate. First-visit uses at most one return for a given pair from each episode. Every-visit can use several returns from one episode when the pair recurs.

containproducecombineEpisodesexperienceVisits to (s, a)one or more per episodeSelected returnsmethod-dependentEstimate of qπ(s, a)more evidence over time
How are returns from visits collected and combined into an increasingly stable estimate?

Convergence with More Visits

As visits to a state-action pair accumulate, both first-visit and every-visit Monte Carlo methods move toward the true expected action value qπ(s, a). The two methods differ in how they select evidence from an episode, but both use returns observed after visits to the pair. When the number of visits to each state-action pair approaches infinity, the source states that both methods converge quadratically to the true expected values.

MethodWithin-episode evidenceLong-run target
First-visit Monte CarloReturn after the first visit in each episodeqπ(s, a)
Every-visit Monte CarloReturns after all visits in each episodeqπ(s, a)

Practice: Classify the Evidence

EASY

A generated episode visits (s, a) at time 2 and again at time 5. The remainder after time 2 produces Return A, and the remainder after time 5 produces Return B. Which return or returns should be used by first-visit Monte Carlo? Which should be used by every-visit Monte Carlo?

Hints
  • First identify the earliest occurrence of (s, a) in the episode.
  • Then ask whether the method keeps only that occurrence or all occurrences.

What do you think happens?

Which returns are selected for the generated episode: first-visit uses Return A or Returns A and B, while every-visit uses Return A or Returns A and B?

  • First-visit uses Return A; every-visit uses Returns A and B
  • First-visit uses Returns A and B; every-visit uses Return A
  • Both methods use only Return A
  • Both methods use Returns A and B
Reveal answer

Answer: First-visit uses Return A; every-visit uses Returns A and B

Time 2 is the first visit in the episode, so first-visit keeps its following return. Every-visit keeps the returns following both occurrences.

  • Counting every occurrence of state s as a visit to (s, a).

    A visit requires both the state and the specified action.

    Fix: Record a visit only when the agent is in s and takes a there.

  • Using only the first return for every-visit estimation.

    Every-visit Monte Carlo includes returns following all visits within the episode.

    Fix: Retain the return after each occurrence of the pair.

  • Using all repeated returns for first-visit estimation.

    First-visit Monte Carlo includes only the return following the first visit within the episode.

    Fix: Keep only the return associated with the earliest occurrence in that episode.

  • Assuming the two methods estimate different action values.

    Both methods target qπ(s, a), the expected return after taking a in s and then following policy π.

    Fix: Distinguish the methods by their return-selection rule, not by their target.

Summary

  1. qπ(s, a) is the expected return after taking action a in state s and then following policy π.
  2. A visit to (s, a) requires encountering state s and taking action a there.
  3. First-visit Monte Carlo uses the return after the first visit to the pair in each episode.
  4. Every-visit Monte Carlo uses the returns after all visits to the pair in each episode.
  5. As visits accumulate, both methods converge to the true expected action values; the source describes this convergence as quadratic.

Key Takeaways

  • Action-value estimation learns qπ(s, a) from returns observed after taking a in s and following policy π.
  • Only a state-action occurrence counts as a visit to the pair.
  • First-visit keeps one return per pair per episode, while every-visit keeps returns after all occurrences.
  • Repeated occurrences can produce distinct returns because each occurrence has a different episode remainder.
  • With infinitely many visits, both methods converge to the true expected action values.