Concepts / Action Preferences and Softmax Action Probabilities

Action Preferences and Softmax Action Probabilities

The gradient bandit algorithm can be understood as a stochastic approximation to gradient ascent.

  • Programming

Reward Feedback Starts the Update

A gradient bandit does not simply count wins and losses for each action. Instead, it maintains action preferences. Those preferences influence the probabilities used to select actions. Stochastic gradient ascent changes the preferences only after two events have occurred: an action has been selected and its reward has been observed.

The update is driven by observed reward feedback, not by action selection alone.

thenafter feedbackAction selectedReward observedPreferences updated
What events must occur before stochastic gradient ascent changes action preferences?

The Moving Reward Baseline

The latest reward is judged relative to a baseline rather than in isolation. That baseline is the average of all rewards received up through and including the current time t. Because it is computed incrementally, the baseline changes as new rewards arrive. It is therefore a moving point of comparison, not a fixed target selected once at the beginning.

ComparisonMeaning for the selected actionDirection of preference change
Current reward above the baselineThe latest outcome was better than the rewards seen so farThe selected action is favored
Current reward below the baselineThe latest outcome was worse than the rewards seen so farThe selected action is discouraged
latest reward is higherlatest reward is lowerReward abovebaselineselected action favoredIncremental rewardbaselineaverage through time tReward belowbaselineselected action discouraged
How does the comparison between the latest reward and the moving baseline determine the update direction?

Reading the Reward Comparison

Suppose the latest reward is better than the running reward average. What qualitative update should you expect?

Compare: Compare the current reward with the average of rewards received through the current time.

Classify: Because the current reward is above the average, the feedback is above-baseline feedback.

Update: Above-baseline feedback favors the selected action and moves the non-selected actions in the opposite direction.

The selected action's future probability is expected to increase qualitatively, while the non-selected actions move oppositely.

Preference Changes Redistribute Probability

Preferences are not the same thing as stored win counts. They are values that influence the probabilities of selecting actions through a softmax action-selection process. When feedback favors the selected action, its preference moves in the favorable direction and the non-selected preferences move oppositely. The resulting softmax probabilities redistribute toward the selected action. When feedback discourages the selected action, the selected action moves in the unfavorable direction and the probability redistribution goes the other way.

preference increases qualitativelyopposite preference movementSelected actionpreference and probabilitybeforeSelected actionfavored afterabove-baseline rewardNon-selectedactionspreferences andprobabilities beforeNon-selected actionsmove oppositely
After reward feedback, how do selected and non-selected preferences change, and how does that affect their softmax probabilities?
influenceproducesAction preferencesSoftmax processAction probabilities
How are action preferences converted into the probabilities used for action selection?

What do you think happens?

The selected action receives a reward below the incremental reward baseline. What happens to its future probability?

  • It increases
  • It decreases
  • It is determined by counting previous wins only
Reveal answer

Answer: It decreases qualitatively.

Below-baseline feedback discourages the selected action. The non-selected actions move oppositely, so the probability distribution shifts away from the selected action.

Selected and Non-Selected Actions

The selected action and the non-selected actions are updated in opposite directions because the feedback is used to favor or discourage the action that produced the observed reward while changing the relative standing of the alternatives. Above-baseline feedback favors the selected action, so the non-selected actions move oppositely. Below-baseline feedback discourages the selected action, so the non-selected actions move in the reverse direction relative to it.

direct preference responseopposite movementchanges relative preferencechanges alternativesReward feedbackSelected actionreward-producing actionProbabilityredistributionNon-selected actionsalternative actions
Why do non-selected actions move in the opposite direction from the selected action after feedback?

When predicting an update, analyze the selected action first: compare its reward with the moving baseline, determine whether it is favored or discouraged, and then infer the opposite movement for the non-selected actions.

The Step-Size Parameter α

The symbol α denotes the positive step-size parameter. It controls the size of a preference adjustment. The reward-baseline comparison determines the qualitative direction: above-baseline feedback favors the selected action, while below-baseline feedback discourages it. α determines how strongly the update is applied. The supplied material establishes the direction and the role of α, but not an exact numerical update amount; that amount also depends on the complete update rule and the action-selection probabilities.

same directionsame directionSmaller αsmaller preference movementUpdate directionset by reward versusbaselineLarger αlarger preference movement
How does changing α alter the update while preserving its direction?

Common Reasoning Errors

  • Treating the latest reward as an absolute judgment

    The reward is judged relative to the incremental average of rewards through the current time, not by its size alone.

    Fix: Compare the current reward with the running reward average before deciding the direction.

  • Thinking preferences are updated when an action is selected

    Stochastic gradient ascent updates preferences after an action has been selected and its reward has been observed.

    Fix: Wait for reward feedback before describing a preference update.

  • Updating only the selected action

    The non-selected actions move oppositely, contributing to the redistribution of action probabilities.

    Fix: Describe both sides of the update: the selected action and the non-selected alternatives.

  • Interpreting the algorithm as a simple win-loss counter

    The gradient bandit maintains preferences that influence action probabilities rather than storing only simple counts of wins and losses.

    Fix: Reason in terms of preferences, reward feedback, the baseline, and softmax probabilities.

  • Ignoring α when discussing update size

    α is the positive step-size parameter, and the exact update amount also depends on the complete update rule and action-selection probabilities.

    Fix: Use α to discuss the scale of the adjustment, while avoiding unsupported numerical values.

Check Your Prediction

MEDIUM

An action is selected and produces a reward below the current incremental reward baseline. Predict the qualitative changes for the selected action, the non-selected actions, and the selected action's future probability. Then explain what α contributes to the update.

Hints
  • Start by comparing the current reward with the moving average.
  • Use the rule for below-baseline feedback.
  • Remember that non-selected actions move oppositely.
  • Describe α as a positive step-size parameter rather than inventing an exact amount.

Practice Solution

An observed reward is below the average of all rewards through the current time. Determine the qualitative update.

Baseline comparison: The current reward is below the incremental reward baseline.

Selected action: Below-baseline feedback discourages the selected action, so its preference moves in the unfavorable direction.

Non-selected actions: The non-selected actions move oppositely to the selected action.

Probability effect: The selected action's future softmax probability decreases qualitatively, while probability shifts away from it toward the alternatives.

Step size: The positive α parameter controls the size of the preference adjustment; the complete rule is needed for an exact amount.

The selected action is discouraged, the non-selected actions move oppositely, and the selected action's future probability decreases qualitatively.

Working Mental Model

  1. Preferences are updated only after an action has been selected and its reward has been observed.
  2. The current reward is compared with an incremental baseline equal to the average of rewards through the current time.
  3. Above-baseline feedback favors the selected action; below-baseline feedback discourages it.
  4. Non-selected actions move oppositely, allowing the softmax action probabilities to redistribute.
  5. α is the positive step-size parameter that controls the magnitude of the preference adjustment, while the complete update rule determines the exact amount.

Key Takeaways

  • A stochastic gradient ascent update begins after action selection and reward observation.
  • The latest reward is judged against a moving incremental reward baseline.
  • Above-baseline feedback increases the selected action's future probability qualitatively; below-baseline feedback decreases it.
  • Non-selected actions move in the opposite direction from the selected action.
  • The positive step-size parameter α controls update magnitude, while the complete rule is required for exact numerical changes.