Action Preferences and Softmax Action Probabilities
The gradient bandit algorithm can be understood as a stochastic approximation to gradient ascent.
Reward Feedback Starts the Update
A gradient bandit does not simply count wins and losses for each action. Instead, it maintains action preferences. Those preferences influence the probabilities used to select actions. Stochastic gradient ascent changes the preferences only after two events have occurred: an action has been selected and its reward has been observed.
The update is driven by observed reward feedback, not by action selection alone.
The Moving Reward Baseline
The latest reward is judged relative to a baseline rather than in isolation. That baseline is the average of all rewards received up through and including the current time t. Because it is computed incrementally, the baseline changes as new rewards arrive. It is therefore a moving point of comparison, not a fixed target selected once at the beginning.
| Comparison | Meaning for the selected action | Direction of preference change |
|---|---|---|
| Current reward above the baseline | The latest outcome was better than the rewards seen so far | The selected action is favored |
| Current reward below the baseline | The latest outcome was worse than the rewards seen so far | The selected action is discouraged |
Reading the Reward Comparison
Suppose the latest reward is better than the running reward average. What qualitative update should you expect?
Compare: Compare the current reward with the average of rewards received through the current time.
Classify: Because the current reward is above the average, the feedback is above-baseline feedback.
Update: Above-baseline feedback favors the selected action and moves the non-selected actions in the opposite direction.
The selected action's future probability is expected to increase qualitatively, while the non-selected actions move oppositely.
Preference Changes Redistribute Probability
Preferences are not the same thing as stored win counts. They are values that influence the probabilities of selecting actions through a softmax action-selection process. When feedback favors the selected action, its preference moves in the favorable direction and the non-selected preferences move oppositely. The resulting softmax probabilities redistribute toward the selected action. When feedback discourages the selected action, the selected action moves in the unfavorable direction and the probability redistribution goes the other way.
What do you think happens?
The selected action receives a reward below the incremental reward baseline. What happens to its future probability?
Reveal answer
Answer: It decreases qualitatively.
Below-baseline feedback discourages the selected action. The non-selected actions move oppositely, so the probability distribution shifts away from the selected action.
Selected and Non-Selected Actions
The selected action and the non-selected actions are updated in opposite directions because the feedback is used to favor or discourage the action that produced the observed reward while changing the relative standing of the alternatives. Above-baseline feedback favors the selected action, so the non-selected actions move oppositely. Below-baseline feedback discourages the selected action, so the non-selected actions move in the reverse direction relative to it.
When predicting an update, analyze the selected action first: compare its reward with the moving baseline, determine whether it is favored or discouraged, and then infer the opposite movement for the non-selected actions.
The Step-Size Parameter α
The symbol α denotes the positive step-size parameter. It controls the size of a preference adjustment. The reward-baseline comparison determines the qualitative direction: above-baseline feedback favors the selected action, while below-baseline feedback discourages it. α determines how strongly the update is applied. The supplied material establishes the direction and the role of α, but not an exact numerical update amount; that amount also depends on the complete update rule and the action-selection probabilities.
Common Reasoning Errors
Treating the latest reward as an absolute judgment
The reward is judged relative to the incremental average of rewards through the current time, not by its size alone.
Fix:
Compare the current reward with the running reward average before deciding the direction.Thinking preferences are updated when an action is selected
Stochastic gradient ascent updates preferences after an action has been selected and its reward has been observed.
Fix:
Wait for reward feedback before describing a preference update.Updating only the selected action
The non-selected actions move oppositely, contributing to the redistribution of action probabilities.
Fix:
Describe both sides of the update: the selected action and the non-selected alternatives.Interpreting the algorithm as a simple win-loss counter
The gradient bandit maintains preferences that influence action probabilities rather than storing only simple counts of wins and losses.
Fix:
Reason in terms of preferences, reward feedback, the baseline, and softmax probabilities.Ignoring α when discussing update size
α is the positive step-size parameter, and the exact update amount also depends on the complete update rule and action-selection probabilities.
Fix:
Use α to discuss the scale of the adjustment, while avoiding unsupported numerical values.
Check Your Prediction
An action is selected and produces a reward below the current incremental reward baseline. Predict the qualitative changes for the selected action, the non-selected actions, and the selected action's future probability. Then explain what α contributes to the update.
Hints
- Start by comparing the current reward with the moving average.
- Use the rule for below-baseline feedback.
- Remember that non-selected actions move oppositely.
- Describe α as a positive step-size parameter rather than inventing an exact amount.
Practice Solution
An observed reward is below the average of all rewards through the current time. Determine the qualitative update.
Baseline comparison: The current reward is below the incremental reward baseline.
Selected action: Below-baseline feedback discourages the selected action, so its preference moves in the unfavorable direction.
Non-selected actions: The non-selected actions move oppositely to the selected action.
Probability effect: The selected action's future softmax probability decreases qualitatively, while probability shifts away from it toward the alternatives.
Step size: The positive α parameter controls the size of the preference adjustment; the complete rule is needed for an exact amount.
The selected action is discouraged, the non-selected actions move oppositely, and the selected action's future probability decreases qualitatively.
Working Mental Model
- Preferences are updated only after an action has been selected and its reward has been observed.
- The current reward is compared with an incremental baseline equal to the average of rewards through the current time.
- Above-baseline feedback favors the selected action; below-baseline feedback discourages it.
- Non-selected actions move oppositely, allowing the softmax action probabilities to redistribute.
- α is the positive step-size parameter that controls the magnitude of the preference adjustment, while the complete update rule determines the exact amount.
Key Takeaways
- A stochastic gradient ascent update begins after action selection and reward observation.
- The latest reward is judged against a moving incremental reward baseline.
- Above-baseline feedback increases the selected action's future probability qualitatively; below-baseline feedback decreases it.
- Non-selected actions move in the opposite direction from the selected action.
- The positive step-size parameter α controls update magnitude, while the complete rule is required for exact numerical changes.