Off-policy Every-Visit Monte Carlo Policy Evaluation
The weighted average update rule incorporates a newly obtained return into an existing estimate.
When a New Return Arrives
In off-policy every-visit Monte Carlo policy evaluation, returns from the same state are combined into an estimate. Each return can have a corresponding random weight. When a new return G_n arrives, the learner must bring the existing estimate up to date instead of treating G_n as an isolated observation or rebuilding the estimate from the beginning.
The new observation is a return-weight pair: the return G_n and its corresponding weight. The previous estimate already represents the earlier returns and weights.
The Two Stored Quantities
For each state, two quantities must be maintained as new returns arrive. V_n is the current weighted-average estimate after the first n returns have been included. C_n is the cumulative sum of the weights assigned to those first n returns. Together, they preserve the information needed to update the estimate when another return is obtained.
Accumulating the Weight History
Suppose the first n−1 returns are G_1 through G_{n−1}, with corresponding weights W_1 through W_{n−1}. Before G_n arrives, the stored estimate summarizes the earlier return-weight pairs, and the stored cumulative value summarizes the earlier weights. When W_n arrives, C_n must become the cumulative sum for the first n weights, not merely W_n by itself.
The subscript n identifies the complete set of the first n processed returns and weights. The update is therefore cumulative: earlier information remains represented while the newest return-weight pair is added.
Combining the New Pair
Tracing G_n, W_n, and V_n
A state has already received returns G_1 through G_{n−1} with corresponding weights W_1 through W_{n−1}. A new return G_n and its corresponding weight W_n now arrive. Trace the information used to produce the updated estimate V_n.
Start with stored information: The stored estimate V_{n−1} represents the earlier returns and weights. The stored cumulative quantity represents the weights for the first n−1 returns.
Receive the new pair: The new observation supplies both G_n and W_n. The return is not processed without its corresponding weight.
Extend the cumulative weight information: The cumulative quantity is updated so that C_n represents the first n weights, including W_n.
Update the estimate: The weighted-average update combines the previous estimate, the new return G_n, the new weight W_n, and the cumulative weight information.
V_n is the current weighted-average estimate after the new return-weight pair has been incorporated.
Repeated Visits to One State
Every-visit evaluation allows returns from the same state to be combined into one estimate. Each visit supplies another return and, in the off-policy setting described here, a corresponding random weight. The estimate is updated as those return-weight pairs arrive, so separate visits contribute to the running estimate rather than being discarded after the first visit.
Mistakes in the Update Trace
Treating G_n as an isolated observation
The method combines the new return with the information already represented by the stored estimate and cumulative sum.
Fix:
Use the previous estimate, the new return, the new return's weight, and the cumulative weight information.Storing only the newest weight
C_n must represent the cumulative sum of the first n weights, not only the newly arrived weight.
Fix:
Carry forward the earlier cumulative information and include W_n.Maintaining only V_n
Both V_n and C_n must be maintained for each state as new returns arrive.
Fix:
Store the estimate and the cumulative weight information together for each state.Rebuilding the estimate from the beginning after every return
The weighted average update rule is incremental and uses the stored estimate and cumulative information.
Fix:
Update the existing estimate when the new return-weight pair arrives.
Practice the State Transition
A state already has an estimate V_{n−1} and cumulative weight information for the first n−1 returns. The next observed pair is G_n and W_n. List the four pieces of information used to produce V_n, and state what C_n must represent after the update.
Hints
- Include the stored estimate.
- Include both parts of the new return-weight pair.
- Remember that the cumulative quantity covers all first n weights.
Practice Solution
Identify the information used when G_n and W_n arrive for a state that already has earlier observations.
Stored estimate: Use V_{n−1}, which represents the earlier weighted estimate.
New observation: Use G_n as the newly obtained return.
New weight: Use W_n as the weight corresponding to G_n.
Cumulative information: Update the cumulative quantity so C_n represents the first n weights.
The update uses V_{n−1}, G_n, W_n, and the cumulative weight information. Afterward, V_n represents the weighted-average estimate after the new return is included.
Key Takeaways
- A new return must be incorporated into the existing weighted estimate rather than handled in isolation.
- Each off-policy return has a corresponding random weight in the setting described by the source.
- C_n records the cumulative sum of the weights for the first n returns.
- For each state, maintain both the current estimate V_n and the cumulative weight information C_n.
- The update from V_{n−1} to V_n combines the previous estimate, the new return G_n, its weight W_n, and the cumulative weight information.
Key Takeaways
- Off-policy every-visit evaluation combines returns from the same state into an estimate.
- A new return arrives with a corresponding weight and must be incorporated into the existing estimate.
- C_n summarizes the cumulative weights for the first n returns.
- Both V_n and C_n must be maintained for each state.
- The incremental update preserves earlier information while adding the newest return-weight pair.