Concepts / Batch Monte Carlo Methods

Batch Monte Carlo Methods

The maximum-likelihood model is built from observed transition frequencies and average observed rewards.

  • Programming

Two Ways to Learn from One Batch

Suppose several episodes have been collected from a Markov process. The same training data can support two different questions. Batch Monte Carlo asks which value estimates fit the returns actually observed in the episodes. Batch TD(0) takes a different route: it first treats the observations as evidence about a Markov process model, then asks which values would be correct if that learned model were the real process.

fitminimizeevaluateconverge towardObserved returnsTraining episodesMean-squared errorOn the training setBatch Monte CarloBest fit to observedreturnsLearned processmodelFrom observed transitionsand rewardsModel-correct valuesCorrect if the model wererealBatch TD(0)Certainty-equivalencetarget
How does batch Monte Carlo update a state from complete observed returns, and how does that differ from batch TD(0), which bootstraps from current value estimates?

From Observations to a Model

The learned model is a maximum-likelihood model built from the batch. Observed transition frequencies provide the model's account of how the process moves between states. Observed rewards are combined into average reward estimates for the relevant states or state-action situations. The result is a model inferred from the data rather than a direct list of the returns seen in individual episodes.

countaverageinforminformcompute values fromCollected episodesObserved transitions andrewardsTransitionfrequenciesObserved movement betweenstatesMaximum-likelihoodmodelModel inferred from thebatchCertainty-equivalenceestimateValues correct if the modelis correctAverage rewardsObserved reward averages
How do observed transitions and rewards become a learned model, and how is that model then used to compute the value estimate as if it were the true environment?

The certainty-equivalence estimate is the value function that would be exactly correct if the learned model were exactly correct. It is called certainty-equivalence because the calculation treats the learned model as though it were the true process.

What the Batch Contains

Batch Monte Carlo keeps its attention on complete returns in the training episodes. If a state appears in several episodes, each appearance can contribute the return observed after that occurrence. The batch estimate is then chosen to minimize mean-squared error on those observed returns. This makes the observed training set, rather than an inferred process model, the direct target.

observed inobserved incollectfitState SAppears in the batchReturn 1Episode 1Batch Monte CarloestimateMinimizes training-seterrorReturn 2Episode 2Observed returnsAll occurrences of S
How do multiple episodes contribute complete returns to the same state, and how are those returns combined into one batch value estimate?

Three Observed Returns

A state appears in three training episodes, producing complete returns of 4, 2, and 6. What estimate best represents the batch Monte Carlo target for this state?

Collect the returns: The state contributes the complete observed returns 4, 2, and 6 to the training set.

Compare candidate estimates: Batch Monte Carlo evaluates candidate values by how well they fit these observed returns, using mean-squared error on the training set.

Choose the best fit: For these three equally represented observations, the estimate that minimizes squared error is their average, which is 4.

The generated example's batch Monte Carlo estimate is 4. The important point is that the target is determined by the observed complete returns, not by first constructing a process model.

Model Correctness versus Data Fit

Batch TD(0) and batch Monte Carlo can use the same collected episodes while defining correctness differently. Batch TD(0) converges to the certainty-equivalence estimate: the values that are correct for the maximum-likelihood model inferred from the episodes. Batch Monte Carlo instead minimizes mean-squared error on the returns in the training episodes.

MethodPrimary evidenceTargetMeaning of correctness
Batch Monte CarloComplete returns in the training episodesValues that minimize mean-squared error on the training setBest fit to the observed returns
Batch TD(0)Maximum-likelihood model inferred from observed transitions and rewardsCertainty-equivalence estimateValues that would be correct if the learned model were the real process

Why Direct Calculation Becomes Difficult

Computing the certainty-equivalence estimate directly requires constructing the maximum-likelihood model and carrying out the full calculation of values for that model. For large state spaces, explicitly storing and processing the complete model and its value calculation can be impractical. TD methods offer a different computational tradeoff: they can approximate the certainty-equivalence solution with memory no more than N and repeated computations over the training set, rather than explicitly carrying out the full direct calculation.

representexpandfeedcan becomemotivateState collectionSmaller state spaceLearned modelExplicit modelrepresentationValue calculationFull direct computationImpractical scaleLarge-state-spacedifficultyState collectionLarge state spaceTD approximationAlternative computationaltradeoff
What grows as the number of states increases when storing the learned model and solving for its value function, and where does the computational bottleneck occur?

When comparing batch methods, state both the target and the computational route. Batch TD(0) is associated with the certainty-equivalence target, while its repeated computations over the training set can avoid explicitly carrying out the full direct calculation. Batch Monte Carlo directly fits the returns present in the training data.

Mistakes in Method Comparison

  • Treating batch Monte Carlo and batch TD(0) as methods with the same batch target.

    Batch Monte Carlo minimizes mean-squared error on observed training returns, while batch TD(0) converges to values correct for the maximum-likelihood model.

    Fix: Identify whether the method is fitting observed returns or evaluating the learned process model.

  • Calling the certainty-equivalence estimate an estimate based only on the returns that happened to appear.

    It is the value function that would be exactly correct if the learned model were exactly correct.

    Fix: Connect the estimate to the learned Markov process model built from transition frequencies and average rewards.

  • Assuming direct certainty-equivalence computation is always practical.

    For large state spaces, explicitly carrying out the full direct calculation can be impractical.

    Fix: Recognize the computational tradeoff offered by TD methods, which can approximate the solution without explicitly performing the full direct calculation.

Check Your Understanding

MEDIUM

A batch contains several episodes. You can either fit values directly to the complete returns in those episodes or infer transition frequencies and average rewards to form a maximum-likelihood model. Which route corresponds to batch Monte Carlo, which corresponds to batch TD(0), and what does each method treat as its definition of correctness?

Hints
  • Ask whether the method is fitting observed returns or evaluating a learned model.
  • Use the certainty-equivalence definition for the model-based route.
  • For the direct-fitting route, name the training-set objective.

Key Takeaways

  1. The certainty-equivalence estimate is the value function that would be exactly correct if the learned Markov process model were exactly correct.
  2. The maximum-likelihood model is built from observed transition frequencies and average observed rewards.
  3. Batch Monte Carlo minimizes mean-squared error on complete returns in the training episodes.
  4. Batch TD(0) converges to values that are correct for the maximum-likelihood model inferred from the batch.
  5. Direct certainty-equivalence computation can be impractical for large state spaces, while TD methods provide an alternative computational tradeoff.

Key Takeaways

  • Batch Monte Carlo fits values to complete observed returns.
  • Batch TD(0) targets the certainty-equivalence estimate associated with a learned maximum-likelihood model.
  • Transition frequencies and average rewards are the evidence used to construct that model.
  • The two methods differ in their definition of correctness, not merely in their update procedures.
  • Direct model-based computation can become impractical for large state spaces, motivating TD-based approximation.