Concepts / Sample-Average Method and Nonstationarity

Sample-Average Method and Nonstationarity

Sample averages are challenged when the values they estimate do not remain fixed.

  • Programming

When Yesterday’s Evidence Becomes Stale

A sample-average action-value method treats past rewards as evidence about the value of an action. That is useful when the action’s reward behavior remains stable. But if the underlying action value keeps changing, some of that evidence describes an older version of the problem. The central question is therefore not only whether an estimate uses many observations, but also whether those observations are still relevant.

time passestime passesInitial q*(a)ten action valuesIndependent walkaction values moveLater q*(a)new action values
What changes from one time step to the next when the true value of each action drifts?

The Modified Ten-Armed Testbed

The modified 10-armed testbed creates nonstationarity by starting with initially equal q*(a) values and then allowing the q*(a) value of each action to take its own independent random walk. The ten action values therefore do not remain fixed. Each action’s underlying value can move as the experiment continues, so rewards collected earlier may describe conditions that are no longer current.

choosecontinue the testbedchange action valuesrepeatEqual q*(a) valuesten actionsSelected actionone of ten actionsIndependent randomwalksq*(a) values changeUpdated q*(a) valuesnext time step
How are the ten action values initialized, selected, and changed over repeated steps?

Following one action through the changing testbed

Imagine tracking one action while the modified testbed runs.

Initialization: The action begins with the same q*(a) status as the other actions because the initial action values are equal.

Repeated steps: The action participates in a problem whose action values take independent random walks. Its underlying value can therefore differ later from its initial value.

Meaning of old rewards: A reward observed earlier is evidence about an earlier condition. If the underlying value has moved, that old evidence may no longer describe the current condition accurately.

The testbed makes the weakness of relying heavily on old rewards visible: the target being estimated changes while observations accumulate.

Two Ways to Update Estimates

MethodStep-size settingHow later observations are weightedRole in the experiment
Sample averageα = 1/nThe step-size becomes smaller as more observations of the relevant action accumulateTests a method whose history becomes increasingly influential
Constant step sizeα = 0.1The step-size remains fixedProvides a contrasting update rule that can help with nonstationary problems

For the sample-average method, n represents how many observations of the relevant action have been included. Its setting α = 1/n means that each new observation receives a smaller step-size as n grows. The method consequently gives the accumulated history increasing influence. By contrast, the comparison method uses α = 0.1 at every update. Its step-size does not shrink merely because more observations have been collected.

weightsresponds with fixed step-sizeα = 1/nstep-size shrinksα = 0.1fixed step-sizeAccumulated historyincreasing influenceChanging conditionscontrasting update
How do the estimates produced by α = 1/n and α = 0.1 respond differently after the underlying action values change?

The before-and-after view captures the adaptation problem. As n grows, the sample-average update becomes less responsive to a newly observed reward. That behavior is reasonable when the action’s value is stable, because more data should produce a more settled estimate. In the modified testbed, however, the action value is not stable, so the same persistence can preserve information about outdated conditions.

Reading the Experiment

A useful experiment compares the two methods inside the same changing environment. Use the modified 10-armed testbed for both methods, keep the exploration setting ε = 0.1 consistent, and record the behavior of each method over repeated steps. The comparison should extend beyond 1000 steps when necessary so that the methods are observed over a sufficiently long portion of the changing problem.

run bothapply repeatedlymeasureaverage across runsModified testbedsame environmentTwo methodsα = 1/n and α = 0.1Repeated stepsε = 0.1Recorded behaviorboth methodsAverage performancecurvesmethod comparison
How do repeated runs produce average performance curves for α = 1/n versus α = 0.1?
  1. Create the modified 10-armed testbed with initially equal q*(a) values and independent random walks for the ten action values.
  2. Run the sample-average method with α = 1/n.
  3. Run the constant-step-size method with α = 0.1.
  4. Keep ε = 0.1 consistent for the methods being compared.
  5. Record both methods over repeated steps, extending the run beyond 1000 steps when necessary.
  6. Compare the resulting behavior as the action values continue to change.

What do you think happens?

After many observations of an action, which method is set up to keep using a fixed step-size when the action values change?

  • The sample-average method with α = 1/n
  • The constant-step-size method with α = 0.1
  • Both methods use a fixed step-size
Reveal answer

Answer: The constant-step-size method with α = 0.1

The sample-average step-size depends on n and becomes smaller as more observations accumulate. The comparison method keeps α fixed at 0.1.

Mistakes in Interpreting Adaptation

  • Assuming that more observations are automatically better.

    Older rewards may describe conditions that are no longer current in a nonstationary problem.

    Fix: Ask whether the reward behavior remains stable before interpreting a longer history as an advantage.

  • Confusing α = 1/n with a constant step-size.

    For the sample-average method, n grows as more observations of the relevant action are included, so the step-size becomes smaller.

    Fix: Track n and recognize that α = 1/n changes with the observation count.

  • Changing the test conditions between methods.

    The comparison would mix the effect of the update rule with the effect of a changed environment or setting.

    Fix: Keep the modified testbed and ε = 0.1 setting consistent while recording both methods.

  • Claiming that constant step-sizes always outperform sample averages.

    The illustration has the narrower purpose of showing why sample averages can struggle when action values change.

    Fix: Describe the result as a comparison of adaptation in this changing problem, not as a claim about every setting.

Practice and Takeaway

MEDIUM

Design a comparison for the modified 10-armed testbed. State what remains constant, identify the two step-size settings, and explain what you would record over repeated steps to compare their adaptation.

Hints
  • Use the same nonstationary testbed for both methods.
  • Include ε = 0.1 as a consistent setting.
  • Compare α = 1/n with α = 0.1.
  • Plan to observe the methods beyond 1000 steps when necessary.
  1. Sample averages work from accumulated evidence and are useful when an action’s reward behavior remains stable. In the modified 10-armed testbed, initially equal q*(a) values take independent random walks, so the underlying problem changes over time. With α = 1/n, each new observation receives a smaller step-size as more observations accumulate. With α = 0.1, the step-size remains constant. A controlled experiment uses the same nonstationary testbed and ε = 0.1 setting for both methods, records their behavior over repeated steps, and extends the run beyond 1000 steps when necessary.

Key Takeaways

  • Sample averages can struggle when the value being estimated changes over time.
  • The modified 10-armed testbed creates nonstationarity by giving initially equal action values independent random walks.
  • The sample-average setting α = 1/n shrinks as more observations of an action accumulate.
  • The comparison setting α = 0.1 keeps a constant step-size and can help with changing problems.
  • A fair experiment keeps the testbed and ε = 0.1 setting consistent while recording both methods over repeated steps.