Variance Reduction in Reinforcement Learning
The baseline changes the REINFORCE update by providing a comparison value for the observed outcome.
The Comparison Idea
REINFORCE uses the outcome of an action to determine how the policy should be updated. REINFORCE with Baseline changes the learning signal by adding a comparison value for that outcome. Instead of considering only whether an outcome was large or small in isolation, the update considers how the outcome relates to the baseline.
The baseline changes the variability of individual updates without changing their expected value. This separation is the central idea of variance reduction in REINFORCE.
Example origin: generated. Imagine judging several outcomes against a typical result. An outcome that looks large by itself may be ordinary in a context where results are usually large. A smaller outcome may be impressive in a context where results are usually small. A baseline supplies that context-sensitive comparison.
From Outcome to Comparison
| Method | Learning signal | Role of the baseline |
|---|---|---|
| Original REINFORCE | The observed outcome | The baseline is uniformly zero |
| REINFORCE with Baseline | The observed outcome compared with a baseline | The baseline supplies a comparison value |
The update changes when the outcome is replaced by an advantage-like quantity: the outcome minus the baseline. A positive comparison means that the outcome is above the selected reference value; a negative comparison means that it is below that reference value. The important distinction is not that REINFORCE with Baseline follows a different expected direction. Its expected update is preserved. The distinction is that each sampled update is measured relative to a comparison value.
Why the Average Is Preserved
Subtracting a baseline changes each sampled update, but it does not change the expected value of the update. The baseline contribution cancels out when the updates are considered in expectation. Therefore, the baseline does not redirect the average learning signal; it changes how individual samples fluctuate around that average.
Same Expected Direction, Different Sample Values
Compare an outcome-only signal with a baseline-adjusted signal for two sampled outcomes.
Choose a comparison value: Example origin: generated. Suppose the selected baseline is 10.
Process the first outcome: For an observed outcome of 12, the comparison signal is 12 minus 10, which is 2.
Process the second outcome: For an observed outcome of 8, the comparison signal is 8 minus 10, which is negative 2.
Interpret the change: The baseline has changed the numerical signals for individual samples by centering them around the comparison value. It has not changed the principle that the expected update is preserved.
The baseline changes individual update values while preserving the expected value of the update.
How Variance Changes
Without a baseline, the outcome itself controls the size of the learning signal. If sampled outcomes vary widely, the resulting updates can also vary widely. Subtracting a comparison value can make the signals more centered and can substantially reduce their spread. The baseline therefore gives learning a way to use the same expected direction while making individual updates less variable.
Example origin: generated. Consider two groups of outcomes. In one group, outcomes are 98, 100, and 102. In another, outcomes are 8, 10, and 12. A single baseline of 10 produces signals of 88, 90, and 92 for the first group, while producing signals of negative 2, 0, and 2 for the second group. This illustrates why a baseline that matches the context matters: subtracting a value that is too low for a high-value context does not provide a useful within-context comparison.
A baseline can reduce variance, but the source claim is that it can affect variance substantially, not that every possible baseline automatically produces the best reduction. The quality and context of the comparison value matter.
State-Specific Comparisons
In a Markov decision process, the baseline should vary with the state. The reason is that the overall values of the available actions can differ from one state to another. A single comparison value may be suitable for one state but unsuitable for another.
One Value or One Value per State
Compare a single baseline with a state-dependent baseline when a process visits a high-value state and a low-value state.
Identify the contexts: Example origin: generated. Suppose one visited state has high overall action values and another visited state has low overall action values.
Use one constant baseline: A single baseline assigns the same comparison value to both states. It may be a reasonable summary for the overall experience, but it is not equally suitable for both contexts.
Use state-dependent baselines: A state-dependent baseline assigns a higher comparison value in the high-value state and a lower comparison value in the low-value state.
Interpret the result: Each outcome is compared with the typical value of its own state. This is why state dependence is appropriate in an MDP.
State-dependent baselines provide context-specific comparison values when action values differ across states.
Three Baseline Choices
| Baseline choice | What it assigns | Meaning |
|---|---|---|
| Zero baseline | Zero in every state | The original REINFORCE method |
| Single baseline value | One value across visited states | A common comparison value, such as the average reward seen so far in a bandit setting |
| State-dependent baseline | A comparison value that varies with the state | A context-specific reference when action values differ across states |
Assuming that adding a baseline changes the expected learning direction.
The baseline preserves the expected value of the update.
Fix:
Separate the unchanged expected value from the potentially changed variance of individual updates.Assuming that every baseline reduces variance equally.
The baseline can affect variance substantially, but the comparison value must be appropriate to the context.
Fix:
Ask whether the baseline gives a useful comparison for the outcomes being evaluated.Using one baseline value for every state in an MDP without considering state values.
The overall values of available actions can differ across states.
Fix:
Use a state-dependent baseline when the state context changes the typical action values.Forgetting that zero is itself a baseline.
Original REINFORCE is the special case in which the baseline is uniformly zero.
Fix:
View REINFORCE with Baseline as a strict generalization that includes the zero-baseline case.
Check Your Understanding
A learner says: “If subtracting a baseline changes every sampled update, it must change the expected update too.” Explain why this conclusion is incorrect. Then describe why a state-dependent baseline is preferable to one constant value when action values differ across states.
Hints
- Distinguish an individual sampled update from the expected update.
- State that the baseline contribution cancels in expectation.
- Connect state dependence to different overall action values in different states.
What do you think happens?
Which choice is the original REINFORCE method?
Reveal answer
Answer: A baseline of zero
Original REINFORCE is the special case of REINFORCE with Baseline in which the baseline is uniformly zero.
Key Takeaways
- REINFORCE with Baseline replaces an outcome-only signal with a comparison between the outcome and a baseline.
- Original REINFORCE is the zero-baseline special case.
- The baseline preserves the expected value of the update because its contribution cancels in expectation.
- The baseline can substantially change, and potentially reduce, the variance of individual updates.
- A state-dependent baseline is appropriate in an MDP because overall action values can differ across states.
Key Takeaways
- REINFORCE with Baseline compares an observed outcome with a baseline rather than using the outcome alone.
- A zero baseline gives the original REINFORCE method.
- Adding a baseline preserves the expected update while changing the variability of individual updates.
- A state-dependent baseline provides an appropriate comparison when different states have different overall action values.
- The essential distinction is unchanged expected value versus potentially reduced variance.