Concepts / DP-Style Expected Backups

DP-Style Expected Backups

Baird's counterexample is a concrete demonstration that semi-gradient TD(0) can be unstable.

  • Programming

When Caution Is Not Enough

A small step size often feels reassuring: if each update is tiny, perhaps the learning process will eventually settle. Baird's counterexample challenges that intuition. In the reported setting, semi-gradient TD(0) sends its weights to infinity for every positive step size. This includes step sizes that are made arbitrarily small. The result is not merely slow learning or temporary oscillation. It is divergence.

The central lesson is that a positive step size can be very small without guaranteeing that semi-gradient TD(0) will converge.

Reading Weight Divergence

The quantity to monitor is the weight vector after repeated updates. In a stable process, the weights would remain bounded or move toward a settled behavior. In Baird's counterexample, repeated semi-gradient TD(0) updates instead drive the weight vector outward without bound. Therefore, divergence means that the weights do not settle at a finite, stable collection of values; they diverge to infinity.

semi-gradient TD(0) updatescontinued updatesWeight vectorinitial valuesWeight vectorafter repeated updatesInfinityreported outcome
What happens to the weights after repeated semi-gradient TD(0) updates in the reported counterexample?

Baird's Counterexample

Baird's counterexample is important because it gives a concrete demonstration that semi-gradient TD(0) can be unstable. The example connects the update procedure to the behavior of the weights: applying semi-gradient TD(0) does not merely produce an inconveniently slow approach to a solution. In the reported setting, the weights diverge to infinity.

settingupdatesrepeated behaviorBaird'scounterexampleSemi-gradient TD(0)repeated updatesWeight vectormonitored after updatesInfinityreported instability
How do the counterexample and the semi-gradient update lead to the reported instability?

This counterexample changes how the step size should be interpreted. Reducing the step size changes how much each individual update moves the weights, but the reported result is that every positive step size still leads to divergence. Consequently, reducing the step size can make the outward behavior appear slower without removing the instability described here.

A Smaller Step Size

Interpreting a Smaller Positive Step Size

Suppose you compare two positive step sizes in the reported setting: one ordinary positive value and one value made much smaller. What conclusion is justified about the long-run behavior?

Track the quantity of interest: Monitor the weight vector after repeated updates rather than judging stability from the size of one update.

Compare the step sizes: The smaller positive step size makes the process appear more cautious at each update, but the source reports instability for every positive step size.

State the long-run conclusion: The smaller step size does not provide a stability guarantee. In the reported setting, the weights still diverge to infinity.

A smaller positive step size may alter the apparent pace of the trajectory, but it does not remove the reported divergence.

Expected Backups Under Pressure

A natural question is whether the instability comes only from using a sampled learning backup. The reported result is stronger: the instability remains when a DP-style expected backup is used. In this version, the weight vector is updated in sweeps through the state space. At every state, the method performs a synchronous semi-gradient backup using the DP full-backup target.

expectationfull-backup targetstate updaterepeated sweepsPossible nextstatesall states consideredExpected valuesDP full-backup targetCurrent statesynchronous backupWeight updatesemi-gradientInfinityreported outcome
How does information from possible next states enter the current-state update, and why does that not prevent the reported divergence?
Backup styleInformation usedReported result in the counterexample
Sampled learning backupA sampled transition or next-state outcomeThe source uses this as the contrast with the expected version
DP-style expected backupThe expectation over possible next states, used in the DP full-backup targetThe instability remains; the weights diverge to infinity

The expected backup removes one possible explanation for the instability: that the method fails only because it relies on a single sampled transition. The source reports divergence even when the backup uses the DP-style expected target. Therefore, averaging over possible next states does not automatically make semi-gradient TD(0) stable in this counterexample.

Common Misreadings

  • Treating a very small positive step size as a guarantee of convergence.

    The reported setting diverges for every positive step size, including arbitrarily small positive values.

    Fix: Treat step-size reduction as a change in update scale, not as an automatic stability guarantee.

  • Describing the behavior as merely slow learning.

    The source characterizes the outcome as divergence to infinity, not simply delayed convergence.

    Fix: Distinguish a bounded or settling trajectory from a weight vector that moves outward without bound.

  • Assuming that an expected backup must be stable because it uses more complete information.

    The source reports that the instability remains with a DP-style expected backup.

    Fix: Ask what the reported behavior is for the complete backup procedure; in this counterexample, the weights still diverge to infinity.

  • Claiming that the source provides a complete numerical trajectory for every weight.

    The source establishes the direction of the behavior but does not supply a complete finite sequence of numerical weight values.

    Fix: Describe the outcome qualitatively as divergence to infinity unless numerical values are explicitly provided.

Check Your Interpretation

MEDIUM

Explain in your own words why replacing a sampled learning backup with a DP-style expected backup does not resolve the instability reported in Baird's counterexample.

Hints
  • Identify what information a DP-style expected backup uses.
  • Recall what happens to the weights after repeated updates in the reported setting.
  • Explain why the result cannot be blamed only on sampling a single next state.

What do you think happens?

If the positive step size is reduced while the method remains in the reported setting, does the source support concluding that the weights will become stable?

  • Yes, every sufficiently small positive step size guarantees stability.
  • No, the reported divergence remains for every positive step size.
  • Only if a sampled backup is used.
  • Only if an expected backup is used.
Reveal answer

Answer: No, the reported divergence remains for every positive step size.

The counterexample is specifically important because making the positive step size arbitrarily small does not remove the reported divergence to infinity.

Practical Takeaway

Use Baird's counterexample as a warning against relying on two tempting shortcuts: making the positive step size very small and replacing a sampled backup with an expected one. In the reported setting, neither change prevents semi-gradient TD(0) from sending its weights to infinity. The example is valuable precisely because it demonstrates instability under both cautious step sizes and DP-style expected backups.

Key Takeaways

  • Divergence means that the semi-gradient TD(0) weight vector moves outward without bound rather than settling at finite values.
  • Baird's counterexample is a concrete demonstration that semi-gradient TD(0) can be unstable.
  • The reported weights diverge to infinity for every positive step size, including arbitrarily small ones.
  • Using a DP-style expected backup does not remove the reported instability.
  • Expected backups therefore should not automatically be treated as guarantees of stability.