Concepts / Off-policy Every-Visit Monte Carlo Policy Evaluation

Off-policy Every-Visit Monte Carlo Policy Evaluation

The weighted average update rule incorporates a newly obtained return into an existing estimate.

  • Programming

When a New Return Arrives

In off-policy every-visit Monte Carlo policy evaluation, returns from the same state are combined into an estimate. Each return can have a corresponding random weight. When a new return G_n arrives, the learner must bring the existing estimate up to date instead of treating G_n as an isolated observation or rebuilding the estimate from the beginning.

The new observation is a return-weight pair: the return G_n and its corresponding weight. The previous estimate already represents the earlier returns and weights.

previous informationcontributesweightsV_{n-1}earlier returns and weightsG_nnew returnW_ncorresponding weightV_nupdated weighted estimate
How does the newly observed return G_n change the previous estimate V_{n-1} to produce V_n?

The Two Stored Quantities

For each state, two quantities must be maintained as new returns arrive. V_n is the current weighted-average estimate after the first n returns have been included. C_n is the cumulative sum of the weights assigned to those first n returns. Together, they preserve the information needed to update the estimate when another return is obtained.

maintainsmaintainscontributeaccumulate intoaffectStatesame stateV_ncurrent estimateC_nsum of first n weightsG_1 through G_nreturns from the stateW_1 through W_ncorresponding weights
What values must be stored for each state, and how are the estimate and cumulative weight connected?

Accumulating the Weight History

Suppose the first n−1 returns are G_1 through G_{n−1}, with corresponding weights W_1 through W_{n−1}. Before G_n arrives, the stored estimate summarizes the earlier return-weight pairs, and the stored cumulative value summarizes the earlier weights. When W_n arrives, C_n must become the cumulative sum for the first n weights, not merely W_n by itself.

retain earlier cumulative informationinclude new weightC_{n-1}first n−1 weightsW_nnew weightC_nfirst n weights
How does C_n accumulate importance-sampling weights and provide cumulative weight information for the update?

The subscript n identifies the complete set of the first n processed returns and weights. The update is therefore cumulative: earlier information remains represented while the newest return-weight pair is added.

Combining the New Pair

Tracing G_n, W_n, and V_n

A state has already received returns G_1 through G_{n−1} with corresponding weights W_1 through W_{n−1}. A new return G_n and its corresponding weight W_n now arrive. Trace the information used to produce the updated estimate V_n.

Start with stored information: The stored estimate V_{n−1} represents the earlier returns and weights. The stored cumulative quantity represents the weights for the first n−1 returns.

Receive the new pair: The new observation supplies both G_n and W_n. The return is not processed without its corresponding weight.

Extend the cumulative weight information: The cumulative quantity is updated so that C_n represents the first n weights, including W_n.

Update the estimate: The weighted-average update combines the previous estimate, the new return G_n, the new weight W_n, and the cumulative weight information.

V_n is the current weighted-average estimate after the new return-weight pair has been incorporated.

previous informationnew observationnew contribution weightcumulative contextV_{n-1}earlier weighted estimateV_nupdated estimateG_nnew returnW_nnew weightC_ncumulative weight
How are the previous weighted estimate, the new return G_n, and the new importance-sampling weight combined in the update?

Repeated Visits to One State

Every-visit evaluation allows returns from the same state to be combined into one estimate. Each visit supplies another return and, in the off-policy setting described here, a corresponding random weight. The estimate is updated as those return-weight pairs arrive, so separate visits contribute to the running estimate rather than being discarded after the first visit.

contributescontributescontributesadds weightadds weightadds weightVisit 1G_1, W_1V_ncombined estimateVisit 2G_2, W_2C_ncumulative weightsVisit nG_n, W_n
How do multiple visits to the same state contribute separate returns and weights to its running estimate?

Mistakes in the Update Trace

  • Treating G_n as an isolated observation

    The method combines the new return with the information already represented by the stored estimate and cumulative sum.

    Fix: Use the previous estimate, the new return, the new return's weight, and the cumulative weight information.

  • Storing only the newest weight

    C_n must represent the cumulative sum of the first n weights, not only the newly arrived weight.

    Fix: Carry forward the earlier cumulative information and include W_n.

  • Maintaining only V_n

    Both V_n and C_n must be maintained for each state as new returns arrive.

    Fix: Store the estimate and the cumulative weight information together for each state.

  • Rebuilding the estimate from the beginning after every return

    The weighted average update rule is incremental and uses the stored estimate and cumulative information.

    Fix: Update the existing estimate when the new return-weight pair arrives.

Practice the State Transition

MEDIUM

A state already has an estimate V_{n−1} and cumulative weight information for the first n−1 returns. The next observed pair is G_n and W_n. List the four pieces of information used to produce V_n, and state what C_n must represent after the update.

Hints
  • Include the stored estimate.
  • Include both parts of the new return-weight pair.
  • Remember that the cumulative quantity covers all first n weights.

Practice Solution

Identify the information used when G_n and W_n arrive for a state that already has earlier observations.

Stored estimate: Use V_{n−1}, which represents the earlier weighted estimate.

New observation: Use G_n as the newly obtained return.

New weight: Use W_n as the weight corresponding to G_n.

Cumulative information: Update the cumulative quantity so C_n represents the first n weights.

The update uses V_{n−1}, G_n, W_n, and the cumulative weight information. Afterward, V_n represents the weighted-average estimate after the new return is included.

Key Takeaways

  1. A new return must be incorporated into the existing weighted estimate rather than handled in isolation.
  2. Each off-policy return has a corresponding random weight in the setting described by the source.
  3. C_n records the cumulative sum of the weights for the first n returns.
  4. For each state, maintain both the current estimate V_n and the cumulative weight information C_n.
  5. The update from V_{n−1} to V_n combines the previous estimate, the new return G_n, its weight W_n, and the cumulative weight information.

Key Takeaways

  • Off-policy every-visit evaluation combines returns from the same state into an estimate.
  • A new return arrives with a corresponding weight and must be incorporated into the existing estimate.
  • C_n summarizes the cumulative weights for the first n returns.
  • Both V_n and C_n must be maintained for each state.
  • The incremental update preserves earlier information while adding the newest return-weight pair.