TD Prediction Methods
The value of n is an important performance setting in n-step TD methods.
Why Vary n
In n-step TD methods, n is not merely a label that identifies one fixed algorithm. It is a performance setting that can be changed and evaluated. The central question is therefore not only whether an n-step method works, but which value of n gives the most useful predictions for the task.
A useful comparison must examine multiple values of n rather than assume that one endpoint is automatically best.
The Experimental Setting
The comparison used a 19-state random walk task. For each parameter setting, the experiment varied both n and α, then measured how well the resulting predictions performed. This separates two questions: which method setting is being tested, and how the quality of its predictions is judged.
The important point is that each value of n defines a separate parameter setting to test. The experiment does not treat n-step TD as having one single, predetermined level of performance.
Reading the Performance Measure
The experiment summarized prediction error across all 19 states, episodes, and repetitions. In other words, the reported performance was not based on whether one particular state received a good prediction. It combined prediction differences across the full set of states and across repeated runs of the task.
The experiment repeated the task 100 times and averaged the first 10 episodes. This gives the reported result a broader experimental basis than a single walk or a single episode.
Interpreting One Parameter Setting
Suppose you want to compare one chosen value of n and one chosen value of α in the random-walk experiment. What must the reported performance represent?
Collect predictions: Consider the predictions produced for the 19 states during the experiment.
Include repeated evidence: Use the results from the episodes included in the experiment and from its repeated runs rather than relying on one walk.
Summarize error: Combine the prediction errors across the states, episodes, and repetitions into one performance result for that parameter setting.
The result is a performance summary for the chosen n and α setting, not a judgment based on one state or one episode.
The Intermediate-Value Pattern
The key reported pattern is that intermediate values of n produced the best performance. The material does not identify one universally best numerical value of n. Instead, it shows that performance can depend on where n lies between the two extreme methods in the comparison.
This result matters because it shows that the best reported setting was not necessarily one of the endpoints. An n-step method can provide a useful middle setting between the two extreme methods, and that middle setting performed best in this experiment.
What the Extremes Reveal
The two extreme methods provide reference points for understanding the n-step family. Comparing only those endpoints would miss the possibility that a value between them performs better. The experiment therefore treats n as something to investigate, not as a choice where an endpoint must be preferred in advance.
A careful interpretation is that intermediate n can balance the weaknesses represented by the two endpoints for the task being studied. The source establishes the performance pattern, but it does not establish one universal causal explanation or prescribe one numerical value of n for all tasks.
When applying n-step TD methods, treat n as an experimental performance setting. Compare several values under the same evaluation procedure instead of assuming that the shortest or longest setting will win.
Common Interpretation Errors
Treating n-step TD as one fixed method with one fixed performance level.
The value of n is an important performance setting, so different values can produce different results.
Fix:
Evaluate multiple n values and compare their measured performance.Judging performance from one state.
The experiment summarizes prediction error across all 19 states, episodes, and repetitions.
Fix:
Use the full performance summary rather than one state-level observation.Assuming the longest or shortest method must be best.
Intermediate values of n produced the best reported performance in the experiment.
Fix:
Treat both extremes as comparison points and test the values between them.Claiming that one numerical value of n is universally optimal.
The provided material does not identify one universally best numerical value.
Fix:
State the result within its experimental context: intermediate values performed best in the reported comparison.
Check Your Interpretation
A researcher evaluates several values of n and α on the 19-state random walk task. The researcher reports one performance result for each parameter setting after using the experiment's repeated state, episode, and repetition evidence. Explain why this evaluation is more informative than testing only one value of n on one episode.
Hints
- Identify what changes when n is treated as a parameter.
- List the dimensions included in the performance summary.
- Use the reported intermediate-value pattern in your explanation.
What do you think happens?
If intermediate values of n performed best in the reported experiment, should you conclude that the same numerical value will always be best in every task?
Reveal answer
Answer: No, because the result is tied to the reported task and comparison.
The source reports that intermediate values performed best in the experiment, but it does not identify one universally best numerical value or claim that the same result holds for every task.
Key Takeaways
- n-step TD methods should be evaluated across multiple values of n because n is an important performance setting.
- The 19-state random walk experiment varied both n and α and summarized prediction error across all 19 states, episodes, and repetitions.
- The experiment repeated the task 100 times and averaged the first 10 episodes.
- Intermediate values of n produced the best reported performance.
- The result supports investigating middle settings rather than assuming that one of the two extreme methods must be best.
Key Takeaways
- The value of n is a performance setting, not merely a fixed label for one method.
- Performance was summarized across all 19 states, episodes, and repeated experiments.
- The reported experiment found the best performance for intermediate values of n.
- The result shows why n should be treated as a parameter to investigate rather than assuming an endpoint is best.
- The experiment does not establish one universally optimal numerical value of n.