Incremental Monte Carlo Methods
The weighted average update rule incorporates a newly obtained return into an existing estimate.
From Episode to Estimate
Incremental Monte Carlo prediction updates a state prediction progressively instead of waiting to perform one final calculation after all experience has been collected. The process is organized episode by episode: an episode produces returns, and those returns provide the information used to adjust the prediction.
The quantity being averaged in Monte Carlo prediction is the return, not the immediate reward.
State Before the New Return
Consider one state that has already appeared in earlier experience. The returns G_1 through G_n−1 that began in this state have already been processed. Their information is summarized by two stored quantities: V_n−1, the current weighted-average estimate, and C_n−1, the cumulative weight assigned to the previously processed returns.
C_n records the cumulative sum of the weights for the first n returns. It is not merely the weight attached to the newest return.
Adding G_n Incrementally
When the next return G_n arrives, it is paired with its corresponding weight W_n in the off-policy setting. The learner does not treat G_n as an isolated observation, and it does not rebuild the estimate from every earlier return. Instead, it combines the stored estimate and stored cumulative weight with this one new return-weight pair.
The sequence matters. Before the update, the stored information represents the first n−1 returns and their weights. After the update, the cumulative weight must represent the first n weights, and V_n must represent the weighted-average estimate after G_n has been included.
Tracing One New Return
A state already has an estimate based on the first n−1 weighted returns. A new return G_n and its weight W_n are now available. What information is used to produce the next estimate?
Start with the stored estimate: Use V_n−1 as the summary of the earlier returns and their weights rather than reopening every earlier observation.
Bring in the new pair: Treat G_n and W_n as one additional return-weight pair supplied by the latest experience.
Update the cumulative weight: Extend the stored cumulative weight so that it represents the weights for the first n returns, not only the newest weight.
Produce the next estimate: Combine the previous estimate, the new return, the new return's weight, and the updated cumulative weight to obtain V_n.
The update is incremental because the previous estimate and cumulative weight summarize earlier work. Only the new return-weight pair must be processed at this step.
Why C_n Matters
The cumulative weight sum C_n tells the update how much weighted information has already been incorporated. It accumulates the weights of the first n returns. Therefore, the newest return is combined with an estimate that already summarizes earlier weighted returns, rather than being treated as though it were the only observation.
A common implementation error is to replace C_n−1 with W_n. The cumulative quantity must include the earlier weights as well as the newly arrived weight.
Returns, Not Rewards
Incremental techniques can be used in more than one setting, but the quantity being averaged changes with the method. Earlier incremental techniques may average rewards observed at individual time steps. Monte Carlo methods average returns. A return is the information produced for the state from the episode, and that return is what the Monte Carlo prediction process incorporates.
Implementing the Update Cycle
An incremental implementation follows a repeated cycle. Process an episode, obtain the returns associated with the states encountered, pair each return with its corresponding weight when the off-policy setting supplies one, and update the stored prediction and cumulative weight for the relevant state. The next episode then contributes another update.
- Maintain V for each state as the current prediction.
- Maintain C for each state as the cumulative sum of the associated weights.
- After an episode, identify the returns that provide the episode's prediction information.
- In the off-policy setting, use each return together with its corresponding random weight.
- Update the state information so that the estimate and cumulative weight include the new contribution.
Mistakes to Avoid
Averaging immediate rewards instead of returns
Monte Carlo methods average returns, and the return is the information supplied by the episode for the state.
Fix:
Identify and incorporate the return associated with the state.Discarding the previous estimate
The new return must be combined with the information already summarized by the stored estimate.
Fix:
Use the previous estimate together with the new return and its weight.Resetting the cumulative weight to the newest weight
C_n must represent the cumulative weights for the first n returns, not only the newest contribution.
Fix:
Extend the previous cumulative weight to include the new return's weight.Updating only one state-wide value without tracking state-specific information
The source specifies that V_n and C_n must be maintained for each state as new returns arrive.
Fix:
Maintain the current estimate and cumulative weight separately for each state.
Check Your Understanding
For one state, describe what must happen when a new return G_n arrives in an off-policy incremental Monte Carlo implementation. Your explanation should name the stored quantities, the new information, and the role of the cumulative weight.
Hints
- Begin with the estimate and cumulative weight from the first n−1 returns.
- Identify the new return and its corresponding weight.
- Explain why the cumulative weight after the update must include all n weights.
What do you think happens?
If an implementation stores only the newest return's weight and forgets the cumulative weight from earlier returns, will it still preserve the intended incremental weighted-average information?
Reveal answer
Answer: No, because the cumulative weight must represent all processed returns.
The stored cumulative sum records the weights assigned to the first n returns. Replacing it with only the newest weight loses information about earlier contributions.
Key Takeaways
- Incremental Monte Carlo prediction updates estimates episode by episode as returns become available.
- In off-policy weighted averaging, a new return G_n is combined with its corresponding weight and the information already stored for the state.
- V_n represents the current weighted-average estimate, while C_n records the cumulative weight of the processed returns.
- Both V_n and C_n must be maintained for each state.
- Monte Carlo prediction averages returns rather than immediate rewards.