Performance Evaluation in Reinforcement Learning
The value of n is an important performance setting in n-step TD methods.
Why n Is a Performance Setting
In n-step temporal-difference methods, n is not merely a technical detail. It is a performance setting that can change how useful the resulting predictions are for a task. Because of this, evaluating one n-step method is not enough. The meaningful question is how performance changes across different values of n.
An n-step TD method should be understood as a family of related methods rather than one fixed setting. Changing n changes which member of that family is being evaluated. The 19-state random walk experiment therefore varied n and compared the resulting prediction performance instead of assuming that one endpoint would automatically be best.
Following the Experiment
- Choose a value of n and a value of α.
- Run the method on the 19-state random walk task.
- Measure the resulting prediction error across all 19 states.
- Repeat the experiment 100 times and average the first 10 episodes.
- Compare the summarized performance with the results for other parameter settings.
This sequence separates three ideas that are easy to mix together. The method setting is the chosen value of n. The additional parameter being varied is α. The evaluation procedure is the way prediction error is collected and summarized. A result is meaningful only when these parts are considered together.
How Prediction Error Is Summarized
The experiment does not judge a parameter setting by looking at only one state. Its performance measure summarizes prediction error across all 19 states. The results are also based on repeated experiments: the experiment was repeated 100 times, and the first 10 episodes were averaged. This gives the comparison a broader basis than a single walk or a single episode.
From state predictions to a comparison result
Suppose two parameter settings produce prediction errors across the full set of states in the 19-state random walk. How should their performance be compared?
Collect state-level errors: For each parameter setting, consider the prediction error across all 19 states rather than selecting one favorable state.
Use repeated runs: Repeat the experiment 100 times so that the comparison is not based on one random walk or one episode.
Average the first 10 episodes: Use the average over the first 10 episodes as part of the reported experimental result.
Compare settings: The resulting summarized values allow the different choices of n and α to be compared on overall prediction performance.
The comparison concerns prediction accuracy across the entire 19-state task and repeated experimental runs, not the behavior of one isolated state.
The Intermediate-n Result
The central reported pattern is that an intermediate value of n produced the best performance in the experiment. The result does not identify one universally best numerical value of n. Instead, it shows that performance can depend on where n lies between the two endpoint methods being compared.
This finding matters because it rejects a simple endpoint-only assumption. It is not enough to compare the two extreme methods and choose whichever appears better. An intermediate n can combine the advantages represented by the broader n-step family well enough to produce more accurate predictions for the task and evaluation setting used in the experiment.
Changing the Amount of Experience
The value of n identifies where a method sits within the n-step family. As n changes, the method changes its position between the two endpoint approaches. This is why the experiment treats n as a parameter to investigate rather than as a fixed label attached to one algorithm.
When evaluating an n-step TD method, report which values of n were tested and compare their results under the same task and evaluation procedure. Otherwise, a statement that the method performed well is incomplete because it does not say which member of the n-step family was used.
Mistakes in Interpreting the Result
Treating n-step TD as one fixed algorithmic setting
The source describes n-step TD as a family whose behavior can be examined for different values of n.
Fix:
Always identify the value of n or explain that multiple values were evaluated.Judging performance from one state
The performance measure summarizes prediction error across all 19 states.
Fix:
Interpret the result as an overall comparison across the full state set.Using one episode as if it were the reported experiment
The experiment was repeated 100 times and averaged over the first 10 episodes.
Fix:
Include the repeated-run and episode-averaging procedure when describing the result.Claiming that one numerical n is always best
The provided result reports an intermediate value as best in this experiment but does not identify a universally best numerical n.
Fix:
State the result in its experimental context.
Check Your Interpretation
A report says that an intermediate value of n performed best in the 19-state random walk experiment. What additional details should you look for before interpreting that claim?
Hints
- Ask whether the result refers to all 19 states or only one state.
- Check whether the experiment used repeated runs and episode averaging.
- Check whether α was also varied.
- Do not assume that the same numerical n must be best in every task.
Interpreting an experimental statement
A student says, 'The largest n is the best because it uses the most information.' Is this conclusion supported by the reported experiment?
Identify the comparison: The experiment compares multiple values of n rather than assuming that one endpoint is best.
Apply the reported pattern: The reported best performance came from an intermediate value of n.
Check the scope: The result belongs to the 19-state random walk experiment and its evaluation procedure.
No. The experiment reports that an intermediate value performed best, so the largest n cannot be declared best from this result.
Key Takeaways
- n is a performance setting, so n-step TD methods should be evaluated across multiple values of n.
- The experiment varied n and α on a 19-state random walk task.
- Performance summarized prediction error across all 19 states, repeated experiments, and the first 10 episodes.
- Intermediate values of n produced the best reported performance.
- The result supports investigating intermediate settings rather than assuming that an endpoint is always best.
Key Takeaways
- n-step TD is a family of methods whose performance can change as n changes.
- The 19-state random walk experiment compared n and α using prediction error across all 19 states.
- Repeating the experiment 100 times and averaging the first 10 episodes provided a broader performance basis.
- Intermediate values of n performed best in the reported experiment.
- That result is experimental and does not establish one universally best numerical value of n.