Concepts / Batch Updating

Batch Updating

Batch updating keeps all episodes seen so far and repeatedly presents that growing batch until convergence.

  • Programming

Learning from a Fixed Record

Many learning methods process experience as it arrives. Batch updating uses a different rhythm. It keeps all episodes seen so far, treats them as one growing collection, and repeatedly presents that collection to the learning method until the value function converges. When a new episode arrives, the earlier episodes are not discarded. They remain part of the batch.

retainedretainedaddedpresent repeatedlyEpisode 1retainedGrowing batchall episodes so farRepeated passesuntil convergenceEpisode 2retainedNew episodeadded
How are previously observed episodes retained, replayed, and incorporated when a new episode is added?

One Complete Batch Pass

Batch updating delays modification of the value function until a complete batch has been processed. Begin with an approximate value function. During one pass through the available experience, calculate the increment associated with every time step at which a nonterminal state is visited. Keep the value function fixed while calculating those increments. After the pass is complete, combine the increments and apply the resulting overall change once. The next pass then uses this newly changed value function.

  1. Collect a finite record of experience, such as several episodes or a fixed number of time steps.
  2. Treat every episode in the record as one batch.
  3. Calculate the required increment for each visited nonterminal state while keeping the current value function fixed.
  4. Apply the combined change after the complete batch has been processed.
  5. Present the same batch again using the updated value function.
  6. Continue repeating complete passes until the value function converges.
processcombineinspectnot yetnext passFinite experienceepisodes or time stepsCalculate incrementsvalue function fixedApply overall changeafter complete passConvergedvalue functionPresent batch againuse updated estimates
What is the sequence from collecting a finite set of episodes to updating estimates and stopping at convergence?

What Monte Carlo Remembers

Batch constant-alpha Monte Carlo converges to the sample averages of the actual returns observed after visits to each state. In other words, its batch result reflects the returns stored in the finite training set. This gives Monte Carlo a limited form of optimality: it minimizes mean-squared error from those training-set returns.

A State Visited with Different Observed Returns

A finite batch contains several visits to one state. The recorded returns after those visits are not all the same. What does batch constant-alpha Monte Carlo move toward?

Collect the observed returns: The batch keeps the returns actually observed after visits to the state.

Compare the observations: The method treats the recorded returns as its training-set evidence for that state.

Replay the batch: Repeated processing causes the estimate to move toward the sample average of those observed returns.

Interpret the result: The resulting estimate is optimal for reproducing the training-set returns in the mean-squared-error sense described for Monte Carlo.

Batch constant-alpha Monte Carlo converges toward the sample average of the actual returns observed after visits to that state.

averaged bymay differ fromevaluated bycompared againstTraining-setreturnsobserved after visitsMonte Carlo estimatesample averageTarget valuefunctioncomparison targetRoot mean-squarederrorreported comparison
How can Monte Carlo be optimal relative to observed episodes while still differing from the target value function?

Why TD(0) Can Win the Comparison

The random-walk comparison separates two questions. First, how closely do estimates reproduce the returns stored in the training set? Second, how closely do estimates match the value function used as the target in the random-walk task? Monte Carlo has the special guarantee for the first question. The reported experiment measured the second question and found that batch TD(0) was consistently better according to root mean-squared error against the target value function.

Question being measuredWhat batch Monte Carlo guaranteesWhat the random-walk result reported
Agreement with observed training-set returnsConverges to their sample averages and minimizes mean-squared error from themThis is Monte Carlo's limited form of optimality
Agreement with the target value functionMay differ from the target because the observed returns are only the finite training setBatch TD(0) was consistently better by root mean-squared error
evaluated againstevaluated againstcomparison resultcomparison resultBatch Monte Carlotraining-set returnsTarget value functionrandom-walk comparisonRoot mean-squarederrorhigher in reportedcomparisonBatch TD(0)same finite batchRoot mean-squarederrorconsistently better
How do batch TD(0) and batch Monte Carlo compare when estimates are judged against the target value function?

The Convergence Boundary

Under batch updating, TD(0) has a deterministic convergence result when the step-size parameter alpha is sufficiently small. Starting from the same available experience and applying the same batch procedure leads to one specific answer. The source describes that answer as independent of the particular sufficiently small choice of alpha.

The condition matters. The claim is not that every possible step size produces the same result. It applies only when alpha is sufficiently small. If that boundary is ignored, the deterministic convergence statement is being interpreted more broadly than the source supports.

Batch constant-alpha Monte Carlo also converges deterministically under the stated conditions, but its convergence point differs from the one reached by TD(0). Therefore, deterministic convergence tells you that a method reaches a specific result from a fixed batch; it does not tell you that two different methods reach the same result.

apply TD(0)yesnointerpret cautiouslyFixed batchsame experienceSufficiently smallalphacondition satisfiedSpecific TD(0) resultdeterministic convergenceStep size notsufficiently smallclaim not establishedNo stated guaranteeoutside the condition
Why does repeated TD(0) processing of a fixed batch have a deterministic result only under the sufficiently-small step-size condition?

Common Interpretation Errors

  • Treating batch updating as a one-time pass through old experience.

    Batch updating keeps all episodes seen so far and repeatedly presents the growing batch until convergence.

    Fix: Retain the earlier episodes, add the new episode, and replay the complete batch.

  • Saying that Monte Carlo is simply the best method.

    Those are different evaluation questions.

    Fix: State which target is being used. The random-walk comparison reported batch TD(0) as consistently better by root mean-squared error against the target value function.

  • Assuming all step sizes support the TD(0) convergence claim.

    The stated result requires alpha to be sufficiently small.

    Fix: Include the step-size condition whenever describing deterministic batch TD(0) convergence.

  • Assuming deterministic convergence means TD(0) and Monte Carlo agree.

    The source states that their convergence points differ.

    Fix: Separate the idea of one reproducible convergence point from the identity of that convergence point.

Check Your Understanding

MEDIUM

A fixed collection of episodes is processed repeatedly. You are told that batch constant-alpha Monte Carlo has converged, while batch TD(0) is being evaluated against a target value function. Explain what Monte Carlo's convergence point represents, why that does not settle the target-value comparison, and which method performed better in the reported random-walk comparison.

Hints
  • Start with the returns actually observed after visits to each state.
  • Separate error against the training-set returns from error against the target value function.
  • Recall the reported root mean-squared-error result for batch TD(0).

What do you think happens?

If a new episode is added to a batch, what happens to the earlier episodes?

  • They are discarded.
  • They remain in the batch and are replayed with the new episode.
  • They are used only once more and then removed.
Reveal answer

Answer: They remain in the batch and are replayed with the new episode.

Batch updating treats all episodes seen so far as one growing collection and repeatedly presents that collection until convergence.

Key Takeaways

  1. Batch updating retains every episode seen so far and repeatedly reprocesses the growing batch.
  2. During a batch pass, increments are calculated while the value function remains fixed; the combined change is applied after the pass.
  3. Batch constant-alpha Monte Carlo converges to sample averages of observed returns and is optimal for mean-squared error against those training-set returns.
  4. The random-walk comparison used a different target: root mean-squared error against the target value function, where batch TD(0) performed consistently better.
  5. Batch TD(0) has a deterministic convergence result when alpha is sufficiently small, and its convergence point differs from that of batch constant-alpha Monte Carlo.

Key Takeaways

  • Batch updating keeps all available experience and replays the complete growing batch until convergence.
  • Batch constant-alpha Monte Carlo converges toward sample averages of the returns observed in the training set.
  • Monte Carlo's optimality is limited to error measured against those observed returns.
  • In the reported random-walk comparison, batch TD(0) performed better by root mean-squared error against the target value function.
  • The deterministic batch TD(0) convergence result depends on alpha being sufficiently small, and it does not imply agreement with batch Monte Carlo.