Concepts / Returns and Episodes in Reinforcement Learning

Returns and Episodes in Reinforcement Learning

The methods differ only in how they select returns from episodes with repeated visits.

  • Programming

One Episode, Two Data Rules

Monte Carlo prediction estimates the value of a state under a policy by collecting episodes that follow that policy and using the returns associated with visits to the state. The important question here is what to do when the same state appears more than once in one episode. First-visit Monte Carlo and every-visit Monte Carlo use the same policy and aim to estimate the same state value, vπ(s). They differ only in which visits contribute returns to the average.

First-visit Monte Carlo records one return per episode for a given state. Every-visit Monte Carlo may record several returns from one episode.

later occurrenceS at t=1return 8S at t=3return 4
Which return belongs to each occurrence of the repeated state?

Tracing Visits Through an Episode

Consider an episode in which state S appears at two different time steps. For a generated example, suppose the return associated with the earlier occurrence is 8 and the return associated with the later occurrence is 4. These are two returns associated with visits to S in the same episode. The methods now make different selection decisions: one keeps only the earliest occurrence, while the other keeps both.

selectepisode orderS at t=1return 88recorded returnS at t=3return 4
How does first-visit Monte Carlo identify the first occurrence of S and ignore later occurrences in the same episode?
selectselectepisode orderS at t=1return 88recorded returnS at t=3return 44recorded return
How does every-visit Monte Carlo include a return for each occurrence of S?

First-Visit Monte Carlo

First-visit Monte Carlo examines an episode from its beginning and finds the earliest occurrence of the state being estimated. It records the return associated with that first occurrence and does not record later occurrences of the same state from that episode. Therefore, one episode contributes at most one return for a given state.

  1. Choose the state whose value is being estimated.
  2. Inspect the episode in time order.
  3. Find the earliest visit to that state.
  4. Record the return associated with that earliest visit.
  5. Stop processing that state for this episode.

Selecting One Return

In a generated episode, state S appears at t=1 with return 8 and at t=3 with return 4. Which return does first-visit Monte Carlo record for S?

Find the earliest occurrence: The first occurrence of S is at t=1.

Use its associated return: The return associated with the first occurrence is 8.

Ignore later occurrences in this episode: The occurrence at t=3 does not add another return for first-visit processing.

The episode contributes one return for S: 8.

Every-Visit Monte Carlo

Every-visit Monte Carlo processes every occurrence of the state in the episode. If a state appears several times, the return associated with each occurrence enters the collection of returns used for the estimate. Consequently, one episode may contribute several returns for one state.

  1. Choose the state whose value is being estimated.
  2. Inspect every time step in the episode.
  3. Identify each visit to that state.
  4. Record the return associated with every identified visit.
  5. Continue until all occurrences in the episode have been processed.

Selecting Several Returns

In a generated episode, state S appears at t=1 with return 8 and at t=3 with return 4. Which returns does every-visit Monte Carlo record for S?

Process the first occurrence: The visit at t=1 contributes its associated return, 8.

Continue to the later occurrence: The visit at t=3 is also a visit to S, so it contributes its associated return, 4.

Collect both returns: The episode contributes both returns rather than stopping after the first occurrence.

The episode contributes two returns for S: 8 and 4.

Same Episode, Different Samples

usesusesusesFirst-visitS at t=1: 8S at t=1return 8Every-visitS at t=1: 8; S at t=3: 4S at t=3return 4
Which state visits and returns are included by each method on the same episode?
QuestionFirst-visit Monte CarloEvery-visit Monte Carlo
Which occurrences are processed?The earliest occurrence of the state in the episodeEvery occurrence of the state in the episode
How many returns can one episode contribute for one state?OneSeveral
Generated episode with returns 8 and 4Records 8Records 8 and 4
Convergence conditionThe number of first visits becomes very largeThe total number of visits becomes very large

The distinction is a data-selection rule, not a policy change. Both methods use episodes obtained by following the same policy, and both aim to estimate vπ(s). Their difference is how many returns from each episode enter the average.

More Visits, More Stable Estimates

A small collection of returns can produce an estimate noticeably different from vπ(s). As more relevant visits are collected, the average generally becomes more stable. First-visit Monte Carlo approaches vπ(s) as the number of first visits increases. Every-visit Monte Carlo approaches vπ(s) as the total number of visits increases. Under the stated conditions, both methods converge to the same state value.

more returns collected8average 88, 4, 6average 6
How do the collected returns and their average change as additional episodes and state visits are processed?

For first-visit Monte Carlo, the source describes the standard deviation of estimation error as falling at a rate of 1 / √n, where n is the number of averaged returns. Every-visit Monte Carlo also has quadratic convergence to vπ(s), although its theoretical justification is less direct because repeated returns within one episode are not treated in the same straightforward way.

Counting and Averaging Mistakes

  • Counting every occurrence when applying first-visit Monte Carlo

    First-visit Monte Carlo records one return per episode for a given state and stops after the earliest occurrence.

    Fix: Keep only the return associated with the earliest occurrence, which is 8 in this generated example.

  • Stopping after the first occurrence when applying every-visit Monte Carlo

    Every-visit Monte Carlo continues through all occurrences of the state in the episode.

    Fix: Record the return associated with each occurrence, such as both 8 and 4.

  • Treating the two methods as different policies

    The distinction concerns which returns are selected from the episode, not the policy that generated it.

    Fix: Hold the policy and the episode-generation process conceptually constant, then apply the appropriate return-selection rule.

  • Using the wrong number of returns in the average

    The average must correspond to the returns actually collected under the selected method.

    Fix: For first-visit processing, count the recorded first visits. For every-visit processing, count all recorded visits.

  • Expecting a small sample to equal vπ(s)

    A small number of returns can produce an estimate noticeably different from the target value.

    Fix: Interpret the estimate as becoming more stable as more relevant visits are collected.

first-visit ruleevery-visit ruleEpisodeS at t=1 and t=31 returnfirst visit only2 returnsboth visits
What is the correct denominator and set of returns when a state appears more than once in an episode?

Practice the Selection Rule

EASY

A generated episode contains state S at three time steps. The associated returns are 7 for the first occurrence, 2 for the second occurrence, and 9 for the third occurrence. List the returns contributed by this episode under first-visit Monte Carlo and under every-visit Monte Carlo. Then state how many returns each method contributes.

Hints
  • First identify the earliest occurrence of S.
  • For first-visit processing, stop after that earliest occurrence.
  • For every-visit processing, continue through all three occurrences.

What do you think happens?

For the practice episode, what returns should be recorded?

  • First-visit: 7; every-visit: 7, 2, 9
  • First-visit: 7, 2, 9; every-visit: 7
  • First-visit: 2; every-visit: 7, 9
  • Both methods: 7
Reveal answer

Answer: First-visit: 7; every-visit: 7, 2, 9

First-visit Monte Carlo keeps only the return from the earliest occurrence. Every-visit Monte Carlo records the return from every occurrence of S.

The Essential Distinction

  1. First-visit Monte Carlo records the return from the first occurrence of a state in each episode.
  2. Every-visit Monte Carlo records the return from every occurrence of that state in each episode.
  3. If a state appears only once in an episode, both methods record the same return from that episode.
  4. The methods use the same policy-generated episodes and aim to estimate the same state value, vπ(s).
  5. As the relevant number of visits increases, both methods converge to vπ(s) under the stated conditions.

Key Takeaways

  • First-visit Monte Carlo selects one return per episode for a given state by using the earliest occurrence.
  • Every-visit Monte Carlo selects a return for every occurrence of the state in the episode.
  • Repeated visits make the methods use different samples, even though both follow the same policy and target the same state value.
  • The averaging denominator must match the number of returns actually recorded under the chosen method.
  • Both methods become more stable and converge to vπ(s) as their relevant visit counts increase.