Root Mean-Squared Error
Batch updating keeps all episodes seen so far and repeatedly presents that growing batch until convergence.
Why the Training Set Matters
A method can be optimal for one objective and still perform worse on another. In this comparison, batch Monte Carlo is optimal for reproducing the returns stored in its training set, while the experiment evaluates methods by how closely their estimates match the target value function in the random-walk task. Root mean-squared error exposes that difference: batch TD(0) was consistently better according to the reported comparison.
The central question is not simply which method averages observed returns more faithfully. It is also what those estimates are being compared against.
How Batch Updating Replays Experience
Batch updating keeps every episode observed so far. When a new episode arrives, the earlier episodes are not discarded. Instead, all episodes seen up to that point are treated as one growing batch, and the selected algorithm is presented with that batch repeatedly until the value function converges. In the random-walk experiment, the constant-alpha methods used a sufficiently small alpha for these repeated presentations.
Adding an Episode to the Batch
Suppose a learner has already observed one episode and then observes a second episode. What happens to the first episode during batch updating?
Keep the first episode: The first episode remains part of the stored experience; it is not removed when the second episode arrives.
Form the growing batch: The first and second episodes are treated together as all episodes seen so far.
Replay the batch: The algorithm is presented with this growing collection repeatedly until the value function converges.
Previously observed episodes continue to influence the batch updates.
What Monte Carlo Converges Toward
Batch constant-alpha Monte Carlo converges to the sample average of the actual returns observed after visits to each state. The important word is observed: the estimate is driven by the returns contained in the training set. Repeated batch presentations move the estimate toward the average represented by those stored returns.
A State Visited in Several Episodes
In a generated training set, a particular state is followed by observed returns of 2, 4, and 6. Toward what value does batch constant-alpha Monte Carlo move its estimate for that state?
Collect the returns: The relevant returns are the actual returns observed after visits to the selected state.
Use their sample average: The sample average of 2, 4, and 6 is 4.
Interpret the result: With repeated batch updating, the state estimate moves toward 4, because that is the average represented by these observed training-set returns.
The estimate converges toward 4 for this generated example.
Two Meanings of Optimal
Monte Carlo has a limited form of optimality. Because batch constant-alpha Monte Carlo converges to the sample averages of the observed returns, it minimizes mean-squared error from those same training-set returns. This does not mean that it necessarily gives the smallest error when evaluated against the target value function used by the task.
Reading the Random-Walk Result
The random-walk comparison separates two questions. First, how well do the estimates reproduce the returns stored in the training set? Second, how well do they match the value function used as the comparison target in the task? Monte Carlo has the special guarantee for the first question. However, when performance was measured against the target value function using root mean-squared error, batch TD(0) was consistently better than batch Monte Carlo.
The result is not a contradiction. Monte Carlo can be best at fitting the observed training-set returns while batch TD(0) is better on the separate evaluation against the target value function.
Mistakes About RMSE Results
Assuming that keeping all episodes means older episodes are discarded when a new episode arrives.
Batch updating keeps all episodes seen so far and repeatedly presents the growing collection.
Fix:
Include the earlier episodes whenever the batch is replayed.Interpreting Monte Carlo's optimality as optimality against every possible target.
Its stated optimality concerns mean-squared error from the training-set returns, while the random-walk result is measured against the target value function.
Fix:
Name the reference being used before deciding what optimal means.Treating a better training-set fit and a better task evaluation as the same result.
The experiment separates reproducing stored returns from matching the target value function.
Fix:
Keep the two evaluation questions separate.
Check Your Interpretation
A learner says: Batch Monte Carlo must beat batch TD(0) in the random-walk experiment because Monte Carlo converges to the average of the observed returns. Explain the mistake in this reasoning.
Hints
- Identify what Monte Carlo is optimal for.
- Identify what the reported root mean-squared error compares against.
- Use the random-walk result to complete the explanation.
What do you think happens?
After a new episode arrives during batch updating, what happens to the previously observed episodes?
Reveal answer
Answer: They remain in the growing batch.
Batch updating keeps all episodes seen so far and repeatedly presents that collection until the value function converges.
Key Takeaways
- Batch updating retains every episode observed so far and repeatedly replays the growing batch until convergence.
- Batch constant-alpha Monte Carlo converges to sample averages of the actual returns observed after visits to each state.
- Monte Carlo's optimality is limited to minimizing error from the training-set returns.
- The random-walk experiment evaluates estimates against a target value function, not merely against the stored training returns.
- Batch TD(0) was consistently better than batch Monte Carlo according to the reported root mean-squared error.
Key Takeaways
- Previously observed episodes remain part of every later batch.
- Batch Monte Carlo moves each state estimate toward the average of its observed returns.
- That training-set optimality does not guarantee the best match to the target value function.
- In the random-walk comparison, batch TD(0) achieved better reported root mean-squared error than batch Monte Carlo.