Sample-Average Method and Nonstationarity
Sample averages are challenged when the values they estimate do not remain fixed.
When Yesterday’s Evidence Becomes Stale
A sample-average action-value method treats past rewards as evidence about the value of an action. That is useful when the action’s reward behavior remains stable. But if the underlying action value keeps changing, some of that evidence describes an older version of the problem. The central question is therefore not only whether an estimate uses many observations, but also whether those observations are still relevant.
The Modified Ten-Armed Testbed
The modified 10-armed testbed creates nonstationarity by starting with initially equal q*(a) values and then allowing the q*(a) value of each action to take its own independent random walk. The ten action values therefore do not remain fixed. Each action’s underlying value can move as the experiment continues, so rewards collected earlier may describe conditions that are no longer current.
Following one action through the changing testbed
Imagine tracking one action while the modified testbed runs.
Initialization: The action begins with the same q*(a) status as the other actions because the initial action values are equal.
Repeated steps: The action participates in a problem whose action values take independent random walks. Its underlying value can therefore differ later from its initial value.
Meaning of old rewards: A reward observed earlier is evidence about an earlier condition. If the underlying value has moved, that old evidence may no longer describe the current condition accurately.
The testbed makes the weakness of relying heavily on old rewards visible: the target being estimated changes while observations accumulate.
Two Ways to Update Estimates
| Method | Step-size setting | How later observations are weighted | Role in the experiment |
|---|---|---|---|
| Sample average | α = 1/n | The step-size becomes smaller as more observations of the relevant action accumulate | Tests a method whose history becomes increasingly influential |
| Constant step size | α = 0.1 | The step-size remains fixed | Provides a contrasting update rule that can help with nonstationary problems |
For the sample-average method, n represents how many observations of the relevant action have been included. Its setting α = 1/n means that each new observation receives a smaller step-size as n grows. The method consequently gives the accumulated history increasing influence. By contrast, the comparison method uses α = 0.1 at every update. Its step-size does not shrink merely because more observations have been collected.
The before-and-after view captures the adaptation problem. As n grows, the sample-average update becomes less responsive to a newly observed reward. That behavior is reasonable when the action’s value is stable, because more data should produce a more settled estimate. In the modified testbed, however, the action value is not stable, so the same persistence can preserve information about outdated conditions.
Reading the Experiment
A useful experiment compares the two methods inside the same changing environment. Use the modified 10-armed testbed for both methods, keep the exploration setting ε = 0.1 consistent, and record the behavior of each method over repeated steps. The comparison should extend beyond 1000 steps when necessary so that the methods are observed over a sufficiently long portion of the changing problem.
- Create the modified 10-armed testbed with initially equal q*(a) values and independent random walks for the ten action values.
- Run the sample-average method with α = 1/n.
- Run the constant-step-size method with α = 0.1.
- Keep ε = 0.1 consistent for the methods being compared.
- Record both methods over repeated steps, extending the run beyond 1000 steps when necessary.
- Compare the resulting behavior as the action values continue to change.
What do you think happens?
After many observations of an action, which method is set up to keep using a fixed step-size when the action values change?
Reveal answer
Answer: The constant-step-size method with α = 0.1
The sample-average step-size depends on n and becomes smaller as more observations accumulate. The comparison method keeps α fixed at 0.1.
Mistakes in Interpreting Adaptation
Assuming that more observations are automatically better.
Older rewards may describe conditions that are no longer current in a nonstationary problem.
Fix:
Ask whether the reward behavior remains stable before interpreting a longer history as an advantage.Confusing α = 1/n with a constant step-size.
For the sample-average method, n grows as more observations of the relevant action are included, so the step-size becomes smaller.
Fix:
Track n and recognize that α = 1/n changes with the observation count.Changing the test conditions between methods.
The comparison would mix the effect of the update rule with the effect of a changed environment or setting.
Fix:
Keep the modified testbed and ε = 0.1 setting consistent while recording both methods.Claiming that constant step-sizes always outperform sample averages.
The illustration has the narrower purpose of showing why sample averages can struggle when action values change.
Fix:
Describe the result as a comparison of adaptation in this changing problem, not as a claim about every setting.
Practice and Takeaway
Design a comparison for the modified 10-armed testbed. State what remains constant, identify the two step-size settings, and explain what you would record over repeated steps to compare their adaptation.
Hints
- Use the same nonstationary testbed for both methods.
- Include ε = 0.1 as a consistent setting.
- Compare α = 1/n with α = 0.1.
- Plan to observe the methods beyond 1000 steps when necessary.
- Sample averages work from accumulated evidence and are useful when an action’s reward behavior remains stable. In the modified 10-armed testbed, initially equal q*(a) values take independent random walks, so the underlying problem changes over time. With α = 1/n, each new observation receives a smaller step-size as more observations accumulate. With α = 0.1, the step-size remains constant. A controlled experiment uses the same nonstationary testbed and ε = 0.1 setting for both methods, records their behavior over repeated steps, and extends the run beyond 1000 steps when necessary.
Key Takeaways
- Sample averages can struggle when the value being estimated changes over time.
- The modified 10-armed testbed creates nonstationarity by giving initially equal action values independent random walks.
- The sample-average setting α = 1/n shrinks as more observations of an action accumulate.
- The comparison setting α = 0.1 keeps a constant step-size and can help with changing problems.
- A fair experiment keeps the testbed and ε = 0.1 setting consistent while recording both methods over repeated steps.