Concepts / Learning State-Value Functions Under a Policy

Learning State-Value Functions Under a Policy

Monte Carlo state-value estimation learns from complete observed returns rather than from a single outcome.

  • Programming

From Episodes to Evidence

Suppose an agent follows a fixed policy and repeatedly experiences episodes. You want to estimate how valuable a particular state is, but you do not begin with the value itself. Instead, you observe what happens after the agent reaches the state, record the resulting return, and use the collected returns as evidence.

The state value is an expected return. A recorded return is not the value itself; it is one observation used to estimate that value. Monte Carlo state-value estimation therefore learns from complete observed returns rather than from a single outcome. As more returns are collected after visits to the state, their average should converge to the expected value.

Pairing States with Returns

An episode supplies a sequence of states and a completed outcome. For each visit to a state, the return observed after that visit becomes evidence about the state's value. The important pairing is therefore state occurrence with the return that follows it, not state occurrence with a single immediate outcome.

paired withpaired withpaired withState at time 1return after visitObserved returnevidence for state valueState at time 2return after visitObserved returnevidence for state valueState at time 3return after visitObserved returnevidence for state value
How does a completed episode produce return observations paired with the states used to estimate their values?

A complete episode can contribute observed returns for the states visited during that episode. The estimate for a state is formed from the returns selected for that state under the chosen visit rule.

Recognizing a Visit

A visit occurs whenever the agent reaches the state during an episode. If state s appears once, the episode contains one visit to s. If state s appears twice or more, each occurrence is a visit to s. The occurrence's position in the episode matters because first-visit and every-visit methods select those occurrences differently.

next stepnext stepnext stepepisode endsState aState svisit 1State bState svisit 2Terminal state
As an episode progresses, which time steps count as visits to a particular state?

Choosing the Visit Rule

MethodReturns selected from an episode containing state sPurpose of the selection
First-visit MCOne return: the return after the first occurrence of sLimits each episode to its first visit to s
Every-visit MCA return after every occurrence of sTreats each occurrence of s as evidence
selectskipselectselectState sfirst occurrenceReturn 1includedState sfirst occurrenceReturn 1includedState slater occurrenceReturn 2not includedState slater occurrenceReturn 2included
When a state appears multiple times in one episode, which occurrences contribute a return to the estimate?

First-visit Monte Carlo includes one return per episode containing the state. It deliberately uses only the return after the first occurrence of that state in the episode. Every-visit Monte Carlo includes a return after every occurrence of the state. If the episode returns to the state, both the return after the first occurrence and the return after the later occurrence enter the every-visit average.

Tracing a Repeated State

Selecting Returns for State s

An exercise provides the observed returns 5, 1, 7, and 3 for occurrences of a state across episodes. Which returns are used by first-visit Monte Carlo and which are used by every-visit Monte Carlo?

Apply the first-visit rule: First-visit Monte Carlo keeps one return per episode containing the state: the return after the first occurrence in that episode. In the exercise, the selected returns are 5 and 3.

Apply the every-visit rule: Every-visit Monte Carlo keeps the return after every occurrence of the state. In the exercise, all four observed returns enter the average: 5, 1, 7, and 3.

Compare the evidence sets: The two methods do not differ because the returns are calculated differently. They differ because they select different occurrences of the state before averaging.

First-visit MC uses 5 and 3. Every-visit MC uses 5, 1, 7, and 3.

selectedselectedselectedselectedselectedselectedOccurrence 1return 5First-visit set5 and 3Occurrence 2return 1Every-visit set5, 1, 7, and 3Occurrence 3return 7Occurrence 4return 3
Which return should be averaged for a repeated state under first-visit estimation, and which returns should be averaged under every-visit estimation?

The averaging step depends on the visit rule chosen before the returns are collected. For the source exercise, first-visit estimation averages the selected returns 5 and 3, whereas every-visit estimation averages 5, 1, 7, and 3.

Following the Episode

generatescontainsproduceaverage selected valuesFixed policyEpisodestates and outcomesState visitsone or more occurrencesObserved returnsevidence for valuesState-value estimateaverage of selected returns
How does an episode unfold under a policy, and how does the completed outcome provide return observations for visited states?

The complete process is empirical. A fixed policy generates episodes. Episodes contain visits to states. The outcomes after those visits provide observed returns. The chosen visit rule determines which of those returns are retained, and the retained returns are averaged to estimate the state value.

Mistakes in Return Selection

  • Treating the state value as identical to one observed return.

    The state value is an expected return, while each recorded return is only an observation used to estimate it.

    Fix: Collect the relevant observed returns and use their average as the estimate.

  • Counting an episode only once even when the state appears more than once.

    Each occurrence of s is a visit, and every-visit MC uses a return after every occurrence.

    Fix: Record the occurrences separately, then apply either the first-visit or every-visit rule.

  • Using every occurrence when first-visit MC was selected.

    First-visit MC deliberately limits each episode to the return after the first occurrence of s.

    Fix: Keep only the first occurrence's return from that episode.

  • Using only the first occurrence when every-visit MC was selected.

    Every-visit MC treats each occurrence of s as evidence.

    Fix: Include the return after every occurrence of s.

Practice the Selection

MEDIUM

A fixed policy generates episodes in which a state appears once in one episode and twice in another. For each method, describe which occurrences contribute returns to the estimate: first-visit Monte Carlo and every-visit Monte Carlo.

Hints
  • For first-visit MC, select at most one return from each episode containing the state.
  • For every-visit MC, select the return after each occurrence of the state.
  • Do not decide from the numerical size of a return; decide from the occurrence selected by the method.

What do you think happens?

If one episode contains two visits to state s, how many returns from that episode can enter the estimate?

  • First-visit MC can use one; every-visit MC can use two.
  • Both methods must use two.
  • First-visit MC must use zero; every-visit MC can use one.
Reveal answer

Answer: First-visit MC can use one; every-visit MC can use two.

First-visit MC includes one return per episode containing the state, while every-visit MC includes a return after every occurrence of the state.

Working Rule

  1. Follow the fixed policy and observe complete episodes.
  2. Identify every occurrence of the state whose value is being estimated.
  3. Record the return observed after each occurrence.
  4. If using first-visit MC, retain only the return after the first occurrence in each episode.
  5. If using every-visit MC, retain the return after every occurrence.
  6. Average the retained returns to estimate the state's expected return.

The central decision is not which return looks most representative. It is which occurrences the selected Monte Carlo method allows into the estimate. First-visit MC uses one return per episode containing the state. Every-visit MC uses a return after every occurrence. In both cases, the retained observed returns provide evidence for the state's expected return.

Key Takeaways

  • A state value is an expected return, while each recorded return is an observation used to estimate it.
  • Every occurrence of a state within an episode counts as a visit.
  • First-visit Monte Carlo includes one return per episode containing the state: the return after its first occurrence.
  • Every-visit Monte Carlo includes the return after every occurrence of the state.
  • The visit rule determines which observed returns are averaged.