Batch Monte Carlo Methods
The maximum-likelihood model is built from observed transition frequencies and average observed rewards.
Two Ways to Learn from One Batch
Suppose several episodes have been collected from a Markov process. The same training data can support two different questions. Batch Monte Carlo asks which value estimates fit the returns actually observed in the episodes. Batch TD(0) takes a different route: it first treats the observations as evidence about a Markov process model, then asks which values would be correct if that learned model were the real process.
From Observations to a Model
The learned model is a maximum-likelihood model built from the batch. Observed transition frequencies provide the model's account of how the process moves between states. Observed rewards are combined into average reward estimates for the relevant states or state-action situations. The result is a model inferred from the data rather than a direct list of the returns seen in individual episodes.
The certainty-equivalence estimate is the value function that would be exactly correct if the learned model were exactly correct. It is called certainty-equivalence because the calculation treats the learned model as though it were the true process.
What the Batch Contains
Batch Monte Carlo keeps its attention on complete returns in the training episodes. If a state appears in several episodes, each appearance can contribute the return observed after that occurrence. The batch estimate is then chosen to minimize mean-squared error on those observed returns. This makes the observed training set, rather than an inferred process model, the direct target.
Three Observed Returns
A state appears in three training episodes, producing complete returns of 4, 2, and 6. What estimate best represents the batch Monte Carlo target for this state?
Collect the returns: The state contributes the complete observed returns 4, 2, and 6 to the training set.
Compare candidate estimates: Batch Monte Carlo evaluates candidate values by how well they fit these observed returns, using mean-squared error on the training set.
Choose the best fit: For these three equally represented observations, the estimate that minimizes squared error is their average, which is 4.
The generated example's batch Monte Carlo estimate is 4. The important point is that the target is determined by the observed complete returns, not by first constructing a process model.
Model Correctness versus Data Fit
Batch TD(0) and batch Monte Carlo can use the same collected episodes while defining correctness differently. Batch TD(0) converges to the certainty-equivalence estimate: the values that are correct for the maximum-likelihood model inferred from the episodes. Batch Monte Carlo instead minimizes mean-squared error on the returns in the training episodes.
| Method | Primary evidence | Target | Meaning of correctness |
|---|---|---|---|
| Batch Monte Carlo | Complete returns in the training episodes | Values that minimize mean-squared error on the training set | Best fit to the observed returns |
| Batch TD(0) | Maximum-likelihood model inferred from observed transitions and rewards | Certainty-equivalence estimate | Values that would be correct if the learned model were the real process |
Why Direct Calculation Becomes Difficult
Computing the certainty-equivalence estimate directly requires constructing the maximum-likelihood model and carrying out the full calculation of values for that model. For large state spaces, explicitly storing and processing the complete model and its value calculation can be impractical. TD methods offer a different computational tradeoff: they can approximate the certainty-equivalence solution with memory no more than N and repeated computations over the training set, rather than explicitly carrying out the full direct calculation.
When comparing batch methods, state both the target and the computational route. Batch TD(0) is associated with the certainty-equivalence target, while its repeated computations over the training set can avoid explicitly carrying out the full direct calculation. Batch Monte Carlo directly fits the returns present in the training data.
Mistakes in Method Comparison
Treating batch Monte Carlo and batch TD(0) as methods with the same batch target.
Batch Monte Carlo minimizes mean-squared error on observed training returns, while batch TD(0) converges to values correct for the maximum-likelihood model.
Fix:
Identify whether the method is fitting observed returns or evaluating the learned process model.Calling the certainty-equivalence estimate an estimate based only on the returns that happened to appear.
It is the value function that would be exactly correct if the learned model were exactly correct.
Fix:
Connect the estimate to the learned Markov process model built from transition frequencies and average rewards.Assuming direct certainty-equivalence computation is always practical.
For large state spaces, explicitly carrying out the full direct calculation can be impractical.
Fix:
Recognize the computational tradeoff offered by TD methods, which can approximate the solution without explicitly performing the full direct calculation.
Check Your Understanding
A batch contains several episodes. You can either fit values directly to the complete returns in those episodes or infer transition frequencies and average rewards to form a maximum-likelihood model. Which route corresponds to batch Monte Carlo, which corresponds to batch TD(0), and what does each method treat as its definition of correctness?
Hints
- Ask whether the method is fitting observed returns or evaluating a learned model.
- Use the certainty-equivalence definition for the model-based route.
- For the direct-fitting route, name the training-set objective.
Key Takeaways
- The certainty-equivalence estimate is the value function that would be exactly correct if the learned Markov process model were exactly correct.
- The maximum-likelihood model is built from observed transition frequencies and average observed rewards.
- Batch Monte Carlo minimizes mean-squared error on complete returns in the training episodes.
- Batch TD(0) converges to values that are correct for the maximum-likelihood model inferred from the batch.
- Direct certainty-equivalence computation can be impractical for large state spaces, while TD methods provide an alternative computational tradeoff.
Key Takeaways
- Batch Monte Carlo fits values to complete observed returns.
- Batch TD(0) targets the certainty-equivalence estimate associated with a learned maximum-likelihood model.
- Transition frequencies and average rewards are the evidence used to construct that model.
- The two methods differ in their definition of correctness, not merely in their update procedures.
- Direct model-based computation can become impractical for large state spaces, motivating TD-based approximation.