DP-Style Expected Backups
Baird's counterexample is a concrete demonstration that semi-gradient TD(0) can be unstable.
When Caution Is Not Enough
A small step size often feels reassuring: if each update is tiny, perhaps the learning process will eventually settle. Baird's counterexample challenges that intuition. In the reported setting, semi-gradient TD(0) sends its weights to infinity for every positive step size. This includes step sizes that are made arbitrarily small. The result is not merely slow learning or temporary oscillation. It is divergence.
The central lesson is that a positive step size can be very small without guaranteeing that semi-gradient TD(0) will converge.
Reading Weight Divergence
The quantity to monitor is the weight vector after repeated updates. In a stable process, the weights would remain bounded or move toward a settled behavior. In Baird's counterexample, repeated semi-gradient TD(0) updates instead drive the weight vector outward without bound. Therefore, divergence means that the weights do not settle at a finite, stable collection of values; they diverge to infinity.
Baird's Counterexample
Baird's counterexample is important because it gives a concrete demonstration that semi-gradient TD(0) can be unstable. The example connects the update procedure to the behavior of the weights: applying semi-gradient TD(0) does not merely produce an inconveniently slow approach to a solution. In the reported setting, the weights diverge to infinity.
This counterexample changes how the step size should be interpreted. Reducing the step size changes how much each individual update moves the weights, but the reported result is that every positive step size still leads to divergence. Consequently, reducing the step size can make the outward behavior appear slower without removing the instability described here.
A Smaller Step Size
Interpreting a Smaller Positive Step Size
Suppose you compare two positive step sizes in the reported setting: one ordinary positive value and one value made much smaller. What conclusion is justified about the long-run behavior?
Track the quantity of interest: Monitor the weight vector after repeated updates rather than judging stability from the size of one update.
Compare the step sizes: The smaller positive step size makes the process appear more cautious at each update, but the source reports instability for every positive step size.
State the long-run conclusion: The smaller step size does not provide a stability guarantee. In the reported setting, the weights still diverge to infinity.
A smaller positive step size may alter the apparent pace of the trajectory, but it does not remove the reported divergence.
Expected Backups Under Pressure
A natural question is whether the instability comes only from using a sampled learning backup. The reported result is stronger: the instability remains when a DP-style expected backup is used. In this version, the weight vector is updated in sweeps through the state space. At every state, the method performs a synchronous semi-gradient backup using the DP full-backup target.
| Backup style | Information used | Reported result in the counterexample |
|---|---|---|
| Sampled learning backup | A sampled transition or next-state outcome | The source uses this as the contrast with the expected version |
| DP-style expected backup | The expectation over possible next states, used in the DP full-backup target | The instability remains; the weights diverge to infinity |
The expected backup removes one possible explanation for the instability: that the method fails only because it relies on a single sampled transition. The source reports divergence even when the backup uses the DP-style expected target. Therefore, averaging over possible next states does not automatically make semi-gradient TD(0) stable in this counterexample.
Common Misreadings
Treating a very small positive step size as a guarantee of convergence.
The reported setting diverges for every positive step size, including arbitrarily small positive values.
Fix:
Treat step-size reduction as a change in update scale, not as an automatic stability guarantee.Describing the behavior as merely slow learning.
The source characterizes the outcome as divergence to infinity, not simply delayed convergence.
Fix:
Distinguish a bounded or settling trajectory from a weight vector that moves outward without bound.Assuming that an expected backup must be stable because it uses more complete information.
The source reports that the instability remains with a DP-style expected backup.
Fix:
Ask what the reported behavior is for the complete backup procedure; in this counterexample, the weights still diverge to infinity.Claiming that the source provides a complete numerical trajectory for every weight.
The source establishes the direction of the behavior but does not supply a complete finite sequence of numerical weight values.
Fix:
Describe the outcome qualitatively as divergence to infinity unless numerical values are explicitly provided.
Check Your Interpretation
Explain in your own words why replacing a sampled learning backup with a DP-style expected backup does not resolve the instability reported in Baird's counterexample.
Hints
- Identify what information a DP-style expected backup uses.
- Recall what happens to the weights after repeated updates in the reported setting.
- Explain why the result cannot be blamed only on sampling a single next state.
What do you think happens?
If the positive step size is reduced while the method remains in the reported setting, does the source support concluding that the weights will become stable?
Reveal answer
Answer: No, the reported divergence remains for every positive step size.
The counterexample is specifically important because making the positive step size arbitrarily small does not remove the reported divergence to infinity.
Practical Takeaway
Use Baird's counterexample as a warning against relying on two tempting shortcuts: making the positive step size very small and replacing a sampled backup with an expected one. In the reported setting, neither change prevents semi-gradient TD(0) from sending its weights to infinity. The example is valuable precisely because it demonstrates instability under both cautious step sizes and DP-style expected backups.
Key Takeaways
- Divergence means that the semi-gradient TD(0) weight vector moves outward without bound rather than settling at finite values.
- Baird's counterexample is a concrete demonstration that semi-gradient TD(0) can be unstable.
- The reported weights diverge to infinity for every positive step size, including arbitrarily small ones.
- Using a DP-style expected backup does not remove the reported instability.
- Expected backups therefore should not automatically be treated as guarantees of stability.