Concepts / Variance Reduction in Reinforcement Learning

Variance Reduction in Reinforcement Learning

The baseline changes the REINFORCE update by providing a comparison value for the observed outcome.

  • Programming

The Comparison Idea

REINFORCE uses the outcome of an action to determine how the policy should be updated. REINFORCE with Baseline changes the learning signal by adding a comparison value for that outcome. Instead of considering only whether an outcome was large or small in isolation, the update considers how the outcome relates to the baseline.

The baseline changes the variability of individual updates without changing their expected value. This separation is the central idea of variance reduction in REINFORCE.

Example origin: generated. Imagine judging several outcomes against a typical result. An outcome that looks large by itself may be ordinary in a context where results are usually large. A smaller outcome may be impressive in a context where results are usually small. A baseline supplies that context-sensitive comparison.

From Outcome to Comparison

MethodLearning signalRole of the baseline
Original REINFORCEThe observed outcomeThe baseline is uniformly zero
REINFORCE with BaselineThe observed outcome compared with a baselineThe baseline supplies a comparison value

The update changes when the outcome is replaced by an advantage-like quantity: the outcome minus the baseline. A positive comparison means that the outcome is above the selected reference value; a negative comparison means that it is below that reference value. The important distinction is not that REINFORCE with Baseline follows a different expected direction. Its expected update is preserved. The distinction is that each sampled update is measured relative to a comparison value.

determinescombined withsubtracted fromdeterminesObserved outcomeREINFORCE signalObserved outcomecomparison beginsREINFORCE updatebaseline = 0Baselinecomparison valueOutcome minusbaselineadvantage-like signalBaseline updateexpected value preserved
What is added or changed when the observed outcome is replaced by outcome minus baseline?

Why the Average Is Preserved

Subtracting a baseline changes each sampled update, but it does not change the expected value of the update. The baseline contribution cancels out when the updates are considered in expectation. Therefore, the baseline does not redirect the average learning signal; it changes how individual samples fluctuate around that average.

contributessubtractedcontainscancels in expectationaverages toObserved outcomesample-specificBaselinecomparison valueIndividual updateoutcome minus baselineBaseline contributioncancels in expectationExpected updateunchanged
How does the baseline term cancel in expectation even though it changes each sampled update?

Same Expected Direction, Different Sample Values

Compare an outcome-only signal with a baseline-adjusted signal for two sampled outcomes.

Choose a comparison value: Example origin: generated. Suppose the selected baseline is 10.

Process the first outcome: For an observed outcome of 12, the comparison signal is 12 minus 10, which is 2.

Process the second outcome: For an observed outcome of 8, the comparison signal is 8 minus 10, which is negative 2.

Interpret the change: The baseline has changed the numerical signals for individual samples by centering them around the comparison value. It has not changed the principle that the expected update is preserved.

The baseline changes individual update values while preserving the expected value of the update.

How Variance Changes

Without a baseline, the outcome itself controls the size of the learning signal. If sampled outcomes vary widely, the resulting updates can also vary widely. Subtracting a comparison value can make the signals more centered and can substantially reduce their spread. The baseline therefore gives learning a way to use the same expected direction while making individual updates less variable.

contributes tocontributes tocan contribute tocan contribute toLow outcomeoutcome-only signalBelow-baselineoutcomeoutcome minus baselineHigh outcomeoutcome-only signalAbove-baselineoutcomeoutcome minus baselineWide update spreadlarger fluctuationsPotentially reducedspreadless variable updates
What changes in the spread of update values when the observed outcome is replaced by outcome minus baseline?

Example origin: generated. Consider two groups of outcomes. In one group, outcomes are 98, 100, and 102. In another, outcomes are 8, 10, and 12. A single baseline of 10 produces signals of 88, 90, and 92 for the first group, while producing signals of negative 2, 0, and 2 for the second group. This illustrates why a baseline that matches the context matters: subtracting a value that is too low for a high-value context does not provide a useful within-context comparison.

A baseline can reduce variance, but the source claim is that it can affect variance substantially, not that every possible baseline automatically produces the best reduction. The quality and context of the comparison value matter.

State-Specific Comparisons

In a Markov decision process, the baseline should vary with the state. The reason is that the overall values of the available actions can differ from one state to another. A single comparison value may be suitable for one state but unsuitable for another.

selectssupportsselectssupportsHigh-value statehigh baselineHigh baselinestate-specific comparisonWithin-statecomparisonoutcome minus baselineLow-value statelow baselineLow baselinestate-specific comparisonWithin-statecomparisonoutcome minus baseline
How can different states have different typical returns, and how does a state-dependent baseline provide the appropriate comparison value in each state?

One Value or One Value per State

Compare a single baseline with a state-dependent baseline when a process visits a high-value state and a low-value state.

Identify the contexts: Example origin: generated. Suppose one visited state has high overall action values and another visited state has low overall action values.

Use one constant baseline: A single baseline assigns the same comparison value to both states. It may be a reasonable summary for the overall experience, but it is not equally suitable for both contexts.

Use state-dependent baselines: A state-dependent baseline assigns a higher comparison value in the high-value state and a lower comparison value in the low-value state.

Interpret the result: Each outcome is compared with the typical value of its own state. This is why state dependence is appropriate in an MDP.

State-dependent baselines provide context-specific comparison values when action values differ across states.

Three Baseline Choices

Baseline choiceWhat it assignsMeaning
Zero baselineZero in every stateThe original REINFORCE method
Single baseline valueOne value across visited statesA common comparison value, such as the average reward seen so far in a bandit setting
State-dependent baselineA comparison value that varies with the stateA context-specific reference when action values differ across states
  • Assuming that adding a baseline changes the expected learning direction.

    The baseline preserves the expected value of the update.

    Fix: Separate the unchanged expected value from the potentially changed variance of individual updates.

  • Assuming that every baseline reduces variance equally.

    The baseline can affect variance substantially, but the comparison value must be appropriate to the context.

    Fix: Ask whether the baseline gives a useful comparison for the outcomes being evaluated.

  • Using one baseline value for every state in an MDP without considering state values.

    The overall values of available actions can differ across states.

    Fix: Use a state-dependent baseline when the state context changes the typical action values.

  • Forgetting that zero is itself a baseline.

    Original REINFORCE is the special case in which the baseline is uniformly zero.

    Fix: View REINFORCE with Baseline as a strict generalization that includes the zero-baseline case.

Check Your Understanding

MEDIUM

A learner says: “If subtracting a baseline changes every sampled update, it must change the expected update too.” Explain why this conclusion is incorrect. Then describe why a state-dependent baseline is preferable to one constant value when action values differ across states.

Hints
  • Distinguish an individual sampled update from the expected update.
  • State that the baseline contribution cancels in expectation.
  • Connect state dependence to different overall action values in different states.

What do you think happens?

Which choice is the original REINFORCE method?

  • A baseline of zero
  • A single nonzero baseline
  • A different nonzero baseline for each state
Reveal answer

Answer: A baseline of zero

Original REINFORCE is the special case of REINFORCE with Baseline in which the baseline is uniformly zero.

Key Takeaways

  1. REINFORCE with Baseline replaces an outcome-only signal with a comparison between the outcome and a baseline.
  2. Original REINFORCE is the zero-baseline special case.
  3. The baseline preserves the expected value of the update because its contribution cancels in expectation.
  4. The baseline can substantially change, and potentially reduce, the variance of individual updates.
  5. A state-dependent baseline is appropriate in an MDP because overall action values can differ across states.

Key Takeaways

  • REINFORCE with Baseline compares an observed outcome with a baseline rather than using the outcome alone.
  • A zero baseline gives the original REINFORCE method.
  • Adding a baseline preserves the expected update while changing the variability of individual updates.
  • A state-dependent baseline provides an appropriate comparison when different states have different overall action values.
  • The essential distinction is unchanged expected value versus potentially reduced variance.