On-policy Monte Carlo Policy Evaluation
Off-policy every-visit Monte Carlo evaluation separates data collection from policy evaluation.
The Evaluation Problem
The central question is this: if one policy generates the data, how can that data evaluate a different policy? Off-policy every-visit Monte Carlo evaluation answers by separating data collection from policy evaluation. A behavior policy generates complete episodes, while a target policy is the policy whose action values are being estimated.
Off-policy every-visit Monte Carlo evaluation is an incremental method that repeatedly generates episodes with a behavior policy, processes each episode backward, and uses weighted importance sampling to adjust updates toward a target policy.
Two Policies, One Episode Source
The behavior policy is responsible for generating complete episodes. The target policy is not used to collect those episodes; instead, the episode is processed afterward to estimate the target policy's action values. Importance sampling provides the correction by comparing the target-policy and behavior-policy probabilities associated with the episode data.
The episode's origin and the policy being evaluated are different roles. The behavior policy supplies experience; importance sampling makes that experience useful for estimating the target policy.
The Backward Pass
After a complete episode is available, the algorithm processes its transitions from the final transition back toward the beginning. At each encountered state-action pair, it uses the return G and the current importance weight W to update Q. Only after that Q update does it change W. Consequently, the weight used for a particular update summarizes the policy correction accumulated from the later part of the episode that has already been processed.
- Start with the final transition of the complete episode.
- Read the return G for the current state-action pair.
- Use the current importance weight W while updating Q.
- Update the cumulative weight C for that state-action pair.
- Change W only after the Q-value update.
- Move to the preceding transition and repeat.
- Stop when the required W equals zero.
One State-Action Update
For one encountered state-action pair, the update has three connected parts. The return G is the episode outcome used for learning. The cumulative quantity C records the total importance weight assigned to that state-action pair. The estimate Q is moved toward G using the current weight. Then W is changed so that earlier state-action pairs receive the correction accumulated from the later part of the episode.
Tracing One Visit
An episode is being processed backward, and the current transition contains an encountered state-action pair with return G. Explain what happens to Q, C, and W at this transition.
Read the current return: Use the return G associated with the current state-action pair. This is the observed outcome that the estimate Q will learn from.
Use the current importance weight: The current W represents the policy correction accumulated from the later portion of the episode already processed in the backward pass.
Update Q and C: Use W with the return-based update of Q, and record the assigned weight in C for this state-action pair. Q is moved toward the observed return rather than being replaced without regard to previous estimates.
Change W afterward: Only after the Q update is complete is W changed using the policy-probability correction for the current step. The resulting W is the weight carried to the next earlier transition.
The correct order is G, then the Q and C update using W, then the W update. Reversing the last two steps changes which correction is used for the current state-action estimate.
Reading a Divergent Trace
An episode trace can differ from what you expect under the target policy because the episode was generated by the behavior policy. The important checkpoint is the action at each state-action pair: compare the action that actually appears in the trace with the action the target policy would favor, and compare the corresponding policy probabilities. Importance sampling uses that comparison to correct the episode data rather than treating the behavior-policy episode as if it had been generated directly by the target policy.
If a trace seems surprising, do not immediately assume that the trace is invalid. First locate the action that differs from the target-policy expectation. Then inspect the backward-pass bookkeeping: G, C, Q, W, and the W = 0 stopping condition. A difference in the observed action affects the importance-sampling correction and therefore can change the result of the update.
Mistakes in Trace Checking
Treating the behavior policy as the policy being evaluated.
The method separates data collection from policy evaluation and uses importance sampling to correct the episode data.
Fix:
Label the behavior policy as the episode generator and the target policy as the policy whose Q estimates are being learned.Updating W before updating Q.
The current update must use the correction accumulated from the later portion of the episode. W is changed only after the Q update.
Fix:
For each backward step, process G, update Q and C with the current W, and then change W.Ignoring C while checking the update.
C records the total weight assigned to each state-action pair and is part of the incremental weighted update.
Fix:
Check C and Q together for every encountered visit.Reading the episode only from beginning to end.
The algorithm processes each episode backward, so later transitions determine the correction carried to earlier ones.
Fix:
Start at the final transition and move toward the beginning.Failing to check the stopping condition.
The source identifies the W = 0 condition as a stopping condition for the backward processing.
Fix:
Check W at the required point in the algorithm and stop when the condition is met.
Practice the Backward Check
A trace contains several state-action pairs and was generated by a behavior policy. For one pair in the middle of the trace, explain what you would inspect before changing Q. Then state what must happen to W after that Q update and what condition can stop further backward processing.
Hints
- Start with the return G and the current importance weight W.
- Remember that C records the total weight assigned to the state-action pair.
- The weight update comes after the Q update.
- Check the W = 0 stopping condition while moving backward.
A strong answer should mention the observed action, the return G, the current W, the cumulative quantity C, the Q update, the subsequent importance-weight change, and the stopping condition. This sequence is more reliable than trying to infer the answer from the final Q value alone.
Convergence and Scope
The source states that Q converges to qπ for all encountered state-action pairs, even when actions are selected according to a potentially different behavior policy. The claim is specifically about state-action pairs that are encountered by the evaluation process. It does not say that an unencountered pair receives an estimate from no data.
- A behavior policy generates complete episodes, while a target policy is evaluated.
- Importance sampling compares target-policy and behavior-policy probabilities to correct the episode data.
- The episode is processed backward, and the current Q update occurs before W changes.
- C records cumulative weight, while Q is moved toward the observed return G.
- Q converges to qπ for all encountered state-action pairs, even when the behavior policy differs from the target policy.
Key Takeaways
- Off-policy every-visit Monte Carlo evaluation separates episode generation from policy evaluation.
- Behavior-policy episodes can evaluate a target policy because importance sampling corrects the data using policy probabilities.
- Backward processing uses G, C, Q, and W in a specific order: update Q and C with the current W, then change W.
- A surprising trace should be checked at the observed action and the resulting importance-weight correction.
- For encountered state-action pairs, Q converges to qπ according to the source's convergence claim.