Batch Updating
Batch updating keeps all episodes seen so far and repeatedly presents that growing batch until convergence.
Learning from a Fixed Record
Many learning methods process experience as it arrives. Batch updating uses a different rhythm. It keeps all episodes seen so far, treats them as one growing collection, and repeatedly presents that collection to the learning method until the value function converges. When a new episode arrives, the earlier episodes are not discarded. They remain part of the batch.
One Complete Batch Pass
Batch updating delays modification of the value function until a complete batch has been processed. Begin with an approximate value function. During one pass through the available experience, calculate the increment associated with every time step at which a nonterminal state is visited. Keep the value function fixed while calculating those increments. After the pass is complete, combine the increments and apply the resulting overall change once. The next pass then uses this newly changed value function.
- Collect a finite record of experience, such as several episodes or a fixed number of time steps.
- Treat every episode in the record as one batch.
- Calculate the required increment for each visited nonterminal state while keeping the current value function fixed.
- Apply the combined change after the complete batch has been processed.
- Present the same batch again using the updated value function.
- Continue repeating complete passes until the value function converges.
What Monte Carlo Remembers
Batch constant-alpha Monte Carlo converges to the sample averages of the actual returns observed after visits to each state. In other words, its batch result reflects the returns stored in the finite training set. This gives Monte Carlo a limited form of optimality: it minimizes mean-squared error from those training-set returns.
A State Visited with Different Observed Returns
A finite batch contains several visits to one state. The recorded returns after those visits are not all the same. What does batch constant-alpha Monte Carlo move toward?
Collect the observed returns: The batch keeps the returns actually observed after visits to the state.
Compare the observations: The method treats the recorded returns as its training-set evidence for that state.
Replay the batch: Repeated processing causes the estimate to move toward the sample average of those observed returns.
Interpret the result: The resulting estimate is optimal for reproducing the training-set returns in the mean-squared-error sense described for Monte Carlo.
Batch constant-alpha Monte Carlo converges toward the sample average of the actual returns observed after visits to that state.
Why TD(0) Can Win the Comparison
The random-walk comparison separates two questions. First, how closely do estimates reproduce the returns stored in the training set? Second, how closely do estimates match the value function used as the target in the random-walk task? Monte Carlo has the special guarantee for the first question. The reported experiment measured the second question and found that batch TD(0) was consistently better according to root mean-squared error against the target value function.
| Question being measured | What batch Monte Carlo guarantees | What the random-walk result reported |
|---|---|---|
| Agreement with observed training-set returns | Converges to their sample averages and minimizes mean-squared error from them | This is Monte Carlo's limited form of optimality |
| Agreement with the target value function | May differ from the target because the observed returns are only the finite training set | Batch TD(0) was consistently better by root mean-squared error |
The Convergence Boundary
Under batch updating, TD(0) has a deterministic convergence result when the step-size parameter alpha is sufficiently small. Starting from the same available experience and applying the same batch procedure leads to one specific answer. The source describes that answer as independent of the particular sufficiently small choice of alpha.
The condition matters. The claim is not that every possible step size produces the same result. It applies only when alpha is sufficiently small. If that boundary is ignored, the deterministic convergence statement is being interpreted more broadly than the source supports.
Batch constant-alpha Monte Carlo also converges deterministically under the stated conditions, but its convergence point differs from the one reached by TD(0). Therefore, deterministic convergence tells you that a method reaches a specific result from a fixed batch; it does not tell you that two different methods reach the same result.
Common Interpretation Errors
Treating batch updating as a one-time pass through old experience.
Batch updating keeps all episodes seen so far and repeatedly presents the growing batch until convergence.
Fix:
Retain the earlier episodes, add the new episode, and replay the complete batch.Saying that Monte Carlo is simply the best method.
Those are different evaluation questions.
Fix:
State which target is being used. The random-walk comparison reported batch TD(0) as consistently better by root mean-squared error against the target value function.Assuming all step sizes support the TD(0) convergence claim.
The stated result requires alpha to be sufficiently small.
Fix:
Include the step-size condition whenever describing deterministic batch TD(0) convergence.Assuming deterministic convergence means TD(0) and Monte Carlo agree.
The source states that their convergence points differ.
Fix:
Separate the idea of one reproducible convergence point from the identity of that convergence point.
Check Your Understanding
A fixed collection of episodes is processed repeatedly. You are told that batch constant-alpha Monte Carlo has converged, while batch TD(0) is being evaluated against a target value function. Explain what Monte Carlo's convergence point represents, why that does not settle the target-value comparison, and which method performed better in the reported random-walk comparison.
Hints
- Start with the returns actually observed after visits to each state.
- Separate error against the training-set returns from error against the target value function.
- Recall the reported root mean-squared-error result for batch TD(0).
What do you think happens?
If a new episode is added to a batch, what happens to the earlier episodes?
Reveal answer
Answer: They remain in the batch and are replayed with the new episode.
Batch updating treats all episodes seen so far as one growing collection and repeatedly presents that collection until convergence.
Key Takeaways
- Batch updating retains every episode seen so far and repeatedly reprocesses the growing batch.
- During a batch pass, increments are calculated while the value function remains fixed; the combined change is applied after the pass.
- Batch constant-alpha Monte Carlo converges to sample averages of observed returns and is optimal for mean-squared error against those training-set returns.
- The random-walk comparison used a different target: root mean-squared error against the target value function, where batch TD(0) performed consistently better.
- Batch TD(0) has a deterministic convergence result when alpha is sufficiently small, and its convergence point differs from that of batch constant-alpha Monte Carlo.
Key Takeaways
- Batch updating keeps all available experience and replays the complete growing batch until convergence.
- Batch constant-alpha Monte Carlo converges toward sample averages of the returns observed in the training set.
- Monte Carlo's optimality is limited to error measured against those observed returns.
- In the reported random-walk comparison, batch TD(0) performed better by root mean-squared error against the target value function.
- The deterministic batch TD(0) convergence result depends on alpha being sufficiently small, and it does not imply agreement with batch Monte Carlo.