Concepts / Reward Baselines in Reinforcement Learning

Reward Baselines in Reinforcement Learning

Stochastic gradient ascent updates preferences only after reward feedback arrives.

  • Programming

The Feedback Moment

A gradient bandit does not update an action merely because the action was selected. The update waits for feedback. Two events must occur: the agent selects an action, and the agent observes the reward produced after that selection. Only then can stochastic gradient ascent change the action preferences.

selects an actionproducestriggers updateAgentEnvironmentReward feedbackR_tAction preferencesupdated
What happens next after an agent selects an action and reward feedback arrives?

The Moving Baseline

The latest reward is not judged in isolation. The algorithm compares it with a running reward average. This average is the baseline for judging the latest outcome, and it includes rewards through the current time t. Because it is computed incrementally, the baseline changes as new rewards arrive. It is therefore a moving point of comparison rather than a fixed target selected once at the beginning.

comparecompareR_t is higherR_t is lowerR_tlatest rewardRunning rewardaveragebaselineReward comparisonabove or belowAbove baselineselected action favoredBelow baselineselected action discouraged
How does the comparison between R_t and the running baseline determine the direction of the preference change?

Selected-Action Probability

An above-baseline reward favors the action that was selected. Its preference changes so that its future probability of selection moves upward. A below-baseline reward discourages the selected action, so its future probability moves downward. This is a qualitative direction of change; the source material does not determine the exact numerical amount of the change.

favorsdiscouragesAbove baselineR_tHigher futureprobabilityselected actionBelow baselineR_tLower futureprobabilityselected action
When the reward is above or below the baseline, does the selected action become more or less likely in the future?

What do you think happens?

Suppose the selected action receives a reward above the running reward average. What happens to its future probability?

  • It increases
  • It decreases
  • The source material does not establish a direction
Reveal answer

Answer: It increases.

Above-baseline feedback favors the selected action. The exact size of the increase requires the complete update rule, but the direction is established.

Opposite Movement

The selected action is not the only action whose preference changes. The non-selected actions move in the opposite direction. When feedback is above the baseline, the selected action is favored and the non-selected actions are discouraged. When feedback is below the baseline, the selected action is discouraged and the non-selected actions move oppositely. This is why the update changes the relative preference among actions rather than simply recording whether one action produced a win or loss.

moves upwardmoves oppositelySelected actionpreferencebefore feedbackNon-selectedpreferencesbefore feedbackSelected actionpreferencefavored afterabove-baseline feedbackNon-selectedpreferencesopposite movement
Why do the preferences of non-selected actions move in the opposite direction from the selected action?

Generated example: Imagine two actions, A and B. If A is selected and receives feedback above the running reward average, the update favors A while moving B in the opposite direction. If A instead receives below-baseline feedback, A is discouraged and B moves oppositely. The example illustrates direction only; it does not specify numerical preferences or probabilities.

The Step-Size Parameter

α is the positive step-size parameter. It controls the scale of a preference update. The baseline comparison determines the qualitative direction: above-baseline feedback favors the selected action, while below-baseline feedback discourages it. α is part of determining how much the preferences change, not whether the feedback is classified as above or below the baseline.

scalesscalesSmaller αpositive step sizeLarger αpositive step sizeSmaller updatesame directionLarger updatesame direction
How does changing α alter the size of each preference update while keeping its direction the same?

What R_t Represents

R_t denotes the reward signal at time t. A reward signal helps define the reinforcement learning problem together with the environment. In this use, R_t is a theoretical abstraction used to describe the feedback available to an agent.

Reward signals and reinforcement signals are distinct concepts in reinforcement learning theory. The model can use the symbol R_t to represent a reward signal at a particular time, while the broader biological system may involve signals with different physical origins and functions. Neural signals are physiological events that may behave like theoretical signals in function, but the theoretical description and the physiological event are not automatically the same thing.

helps definemay relate functionally tomay involvedoes not identify one literal signalR_ttheoretical reward signalReinforcementlearning problemdefined with environmentReinforcementsignaldistinct conceptPhysiological neuralsignalsbiological eventsMany neural signalsand systemsbiological complexity
How are R_t and reinforcement signals related to physiological signals in the brain, and why is R_t not a literal single master reward signal in an animal?

Reasoning Practice

Classifying Two Feedback Outcomes

An agent selects an action and then receives feedback. In one case, the feedback is above the running reward average. In another case, it is below the running reward average. Predict the direction of the selected action's future probability and the movement of the non-selected actions.

Above-baseline case: The selected action is favored, so its future probability moves upward. The non-selected actions move in the opposite direction.

Below-baseline case: The selected action is discouraged, so its future probability moves downward. The non-selected actions move oppositely.

Update timing: Neither preference change occurs immediately after selection alone. The action must first be selected and its reward must then be observed.

Update size: The positive step-size parameter α contributes to the size of the update. The exact numerical amount cannot be determined without the complete update rule and the action-selection probabilities.

Compare R_t with the moving reward average first. Use the comparison for direction, and treat α as a positive parameter involved in the update size.

EASY

An action has just been selected, but no reward has yet been observed. Should stochastic gradient ascent update the action preferences at this moment? Explain what event is still required and what comparison will be made once it arrives.

Hints
  • Recall the two events required before preferences change.
  • The latest reward is compared with a running average rather than a fixed initial target.
MEDIUM

A learner says, R_t must be one physical master reward signal in an animal's brain because the model uses one symbol. Correct the statement using the distinction between theoretical signals and physiological events.

Hints
  • R_t is a theoretical reward signal at time t.
  • Neural signals are physiological events that may behave like theoretical signals in function.
  • One model symbol does not establish one unchanged biological signal.

Common Mistakes

  • Updating preferences as soon as an action is selected.

    Stochastic gradient ascent updates preferences only after both action selection and reward observation.

    Fix: Wait for the reward feedback, then compare it with the running reward average.

  • Comparing the latest reward with a fixed target chosen at the beginning.

    The baseline is the incrementally computed average of rewards through the current time t.

    Fix: Use the moving reward average as the point of comparison.

  • Assuming that a large reward always favors the selected action.

    The relevant question is whether the reward is above or below the rewards seen so far, not whether its isolated numerical value sounds large.

    Fix: Judge R_t relative to the baseline.

  • Forgetting the non-selected actions.

    The non-selected actions move in the opposite direction.

    Fix: Track both the selected action and the non-selected actions.

  • Treating α as the factor that decides whether feedback is good or bad.

    The baseline comparison determines the direction; α is the positive step-size parameter involved in the update size.

    Fix: Separate direction from magnitude.

  • Treating R_t as one literal master reward signal in an animal's brain.

    R_t is a theoretical abstraction, while neural signals are physiological events and biological systems may involve many signals and systems.

    Fix: Keep the model-level signal distinct from claims about a particular physiological event.

Key Takeaways

  1. Preferences change only after an action has been selected and its reward has been observed.
  2. The current reward R_t is judged against an incrementally updated running reward average.
  3. Above-baseline feedback favors the selected action; below-baseline feedback discourages it.
  4. Non-selected actions move in the opposite direction, while α is the positive step-size parameter that contributes to update size.
  5. R_t is a theoretical reward signal, not automatically a single physical master reward signal in an animal's brain.

Key Takeaways

  • Stochastic gradient ascent waits for reward feedback before changing action preferences.
  • The running reward average provides a moving baseline for evaluating R_t.
  • Above-baseline feedback increases the selected action's future probability, while below-baseline feedback decreases it.
  • Non-selected actions move oppositely, and α controls the scale of the preference update as a positive step-size parameter.
  • R_t is a theoretical model signal and should not be identified with one unitary physiological reward signal in an animal's brain.