Concepts / Root Mean-Squared Error

Root Mean-Squared Error

Batch updating keeps all episodes seen so far and repeatedly presents that growing batch until convergence.

  • Programming

Why the Training Set Matters

A method can be optimal for one objective and still perform worse on another. In this comparison, batch Monte Carlo is optimal for reproducing the returns stored in its training set, while the experiment evaluates methods by how closely their estimates match the target value function in the random-walk task. Root mean-squared error exposes that difference: batch TD(0) was consistently better according to the reported comparison.

The central question is not simply which method averages observed returns more faithfully. It is also what those estimates are being compared against.

How Batch Updating Replays Experience

Batch updating keeps every episode observed so far. When a new episode arrives, the earlier episodes are not discarded. Instead, all episodes seen up to that point are treated as one growing batch, and the selected algorithm is presented with that batch repeatedly until the value function converges. In the random-walk experiment, the constant-alpha methods used a sufficiently small alpha for these repeated presentations.

formnew episode arrivesretain earlier episoderepeat until convergenceEpisode 1storedBatch 1present repeatedlyEpisode 2addedBatch 1 + 2present repeatedlyConverged estimatesvalue function
How does the growing collection of previously observed episodes get replayed until the estimates converge?

Adding an Episode to the Batch

Suppose a learner has already observed one episode and then observes a second episode. What happens to the first episode during batch updating?

Keep the first episode: The first episode remains part of the stored experience; it is not removed when the second episode arrives.

Form the growing batch: The first and second episodes are treated together as all episodes seen so far.

Replay the batch: The algorithm is presented with this growing collection repeatedly until the value function converges.

Previously observed episodes continue to influence the batch updates.

What Monte Carlo Converges Toward

Batch constant-alpha Monte Carlo converges to the sample average of the actual returns observed after visits to each state. The important word is observed: the estimate is driven by the returns contained in the training set. Repeated batch presentations move the estimate toward the average represented by those stored returns.

averagerepeated batch updatesObserved returnsafter visits to a stateSample averagetraining-set returnsState estimateafter convergence
How do repeated batch Monte Carlo updates move a state estimate toward the average of the observed returns?

A State Visited in Several Episodes

In a generated training set, a particular state is followed by observed returns of 2, 4, and 6. Toward what value does batch constant-alpha Monte Carlo move its estimate for that state?

Collect the returns: The relevant returns are the actual returns observed after visits to the selected state.

Use their sample average: The sample average of 2, 4, and 6 is 4.

Interpret the result: With repeated batch updating, the state estimate moves toward 4, because that is the average represented by these observed training-set returns.

The estimate converges toward 4 for this generated example.

Two Meanings of Optimal

Monte Carlo has a limited form of optimality. Because batch constant-alpha Monte Carlo converges to the sample averages of the observed returns, it minimizes mean-squared error from those same training-set returns. This does not mean that it necessarily gives the smallest error when evaluated against the target value function used by the task.

converges towardminimizes error fromevaluated againstdetermines reported comparisonBatch Monte Carlosample averagesObserved returnstraining setTarget value functiontask comparisonMinimum training-seterrorlimited optimalityReported task errorRMSE comparison
What value function does batch Monte Carlo converge to, and how does that differ from the value function used as the comparison target?

Reading the Random-Walk Result

The random-walk comparison separates two questions. First, how well do the estimates reproduce the returns stored in the training set? Second, how well do they match the value function used as the comparison target in the task? Monte Carlo has the special guarantee for the first question. However, when performance was measured against the target value function using root mean-squared error, batch TD(0) was consistently better than batch Monte Carlo.

evaluate againstevaluate againstreported RMSE outcomereported RMSE outcomeBatch TD(0)random-walk methodBatch Monte Carlorandom-walk methodTarget value functionevaluation referenceBetter RMSEconsistently reportedHigher RMSErelative comparison
How did the two methods compare when their estimates were measured against the target value function by root mean-squared error?

The result is not a contradiction. Monte Carlo can be best at fitting the observed training-set returns while batch TD(0) is better on the separate evaluation against the target value function.

Mistakes About RMSE Results

  • Assuming that keeping all episodes means older episodes are discarded when a new episode arrives.

    Batch updating keeps all episodes seen so far and repeatedly presents the growing collection.

    Fix: Include the earlier episodes whenever the batch is replayed.

  • Interpreting Monte Carlo's optimality as optimality against every possible target.

    Its stated optimality concerns mean-squared error from the training-set returns, while the random-walk result is measured against the target value function.

    Fix: Name the reference being used before deciding what optimal means.

  • Treating a better training-set fit and a better task evaluation as the same result.

    The experiment separates reproducing stored returns from matching the target value function.

    Fix: Keep the two evaluation questions separate.

Check Your Interpretation

MEDIUM

A learner says: Batch Monte Carlo must beat batch TD(0) in the random-walk experiment because Monte Carlo converges to the average of the observed returns. Explain the mistake in this reasoning.

Hints
  • Identify what Monte Carlo is optimal for.
  • Identify what the reported root mean-squared error compares against.
  • Use the random-walk result to complete the explanation.

What do you think happens?

After a new episode arrives during batch updating, what happens to the previously observed episodes?

  • They are discarded.
  • They remain in the growing batch.
  • They are used only once and then removed.
  • They are replaced by the new episode.
Reveal answer

Answer: They remain in the growing batch.

Batch updating keeps all episodes seen so far and repeatedly presents that collection until the value function converges.

Key Takeaways

  1. Batch updating retains every episode observed so far and repeatedly replays the growing batch until convergence.
  2. Batch constant-alpha Monte Carlo converges to sample averages of the actual returns observed after visits to each state.
  3. Monte Carlo's optimality is limited to minimizing error from the training-set returns.
  4. The random-walk experiment evaluates estimates against a target value function, not merely against the stored training returns.
  5. Batch TD(0) was consistently better than batch Monte Carlo according to the reported root mean-squared error.

Key Takeaways

  • Previously observed episodes remain part of every later batch.
  • Batch Monte Carlo moves each state estimate toward the average of its observed returns.
  • That training-set optimality does not guarantee the best match to the target value function.
  • In the random-walk comparison, batch TD(0) achieved better reported root mean-squared error than batch Monte Carlo.