Soft-max Action Selection
Numerical preferences express how strongly each action is favored.
From Favoring to Choosing
An action-selection method needs a way to express which actions are currently favored. In gradient bandit methods, that role is handled by a numerical preference for every action. The preference indicates how strongly an action is favored, but it is not an estimate of the reward that action will produce.
A preference is a relative favoring signal. The preference itself is not the final decision and is not an estimate of reward.
The Three-Stage Selection Process
Soft-max action selection can be understood as a three-stage process. First, each action a has a numerical preference H_t(a) at time t. Second, the preferences are processed by a soft-max distribution. Third, the resulting probability π_t(a) determines how likely that action is to be taken at time t. A larger preference makes an action more strongly favored relative to actions with smaller preferences.
The soft-max distribution is therefore the link between an internal preference state and an action choice. It does not simply select the action with the largest preference as a fixed rule. Instead, it maps the collection of preferences to probabilities, so the relative preferences determine how strongly each action is favored.
Equal Preferences at the Start
Three actions begin equally favored
Suppose three actions begin with equal numerical preferences.
Record the preferences: Each action has the same initial preference, so no action is relatively more favored than another.
Apply soft-max: The soft-max distribution receives equal preferences and therefore assigns equal initial probabilities to the actions.
Interpret the result: The actions begin equally likely to be selected. This is a starting condition, not a guarantee that their probabilities will remain equal.
Equal initial preferences produce equal initial action-selection probabilities.
Why Absolute Values Do Not Matter
Soft-max action selection depends on relative preferences rather than on the absolute numerical labels attached to them. If one action has a larger preference than another, it is more strongly favored relative to that other action. What matters is this relationship between preferences.
For example, adding 1000 to every preference changes the numerical labels but does not change the relative preferences. Because soft-max selection depends on those relative preferences, the action probabilities remain unchanged. This is why a preference value should be interpreted in comparison with the other preferences, not in isolation.
Mistakes About Preference Values
Treating H_t(a) as an estimate of reward
In gradient bandit methods, the preference indicates how strongly an action is favored, but it is not itself an estimate of reward.
Fix:
Interpret H_t(a) as a preference signal that soft-max converts into an action probability.Thinking the largest preference is the entire selection rule
The preferences are processed by a soft-max distribution, which produces probabilities for taking actions.
Fix:
Follow the full process: preferences are transformed into probabilities, and those probabilities determine how likely each action is to be taken.Believing that equal initial probabilities remain equal
The initial equality only reflects equal initial preferences. Later relative changes can produce new action probabilities.
Fix:
Reconsider the relative preferences whenever the preference values change.Assuming that adding the same constant changes the probabilities
The common shift changes absolute values but leaves relative preferences unchanged.
Fix:
Compare the preferences with one another; a common constant does not affect the resulting action probabilities.
Check Your Understanding
Imagine that every action in a selection problem receives the same increase in preference. Will the soft-max action probabilities change? Explain your answer using the idea of relative preferences.
Hints
- Compare each action's preference with the other actions' preferences before and after the increase.
- Ask whether the ordering or differences in relative favoring have changed.
What do you think happens?
Three actions begin with equal preferences. Before reading the reveal, predict whether their initial soft-max probabilities are equal or unequal.
Reveal answer
Answer: Equal
Equal initial preferences mean that no action is relatively more favored than another, so soft-max produces equal initial action-selection probabilities.
Key Takeaways
- H_t(a) is the numerical preference for action a at time t.
- A preference expresses how strongly an action is favored; it is not itself an estimate of reward.
- Soft-max transforms the collection of preferences into action probabilities π_t(a).
- Relative preferences determine the probabilities, so adding the same constant to every preference leaves them unchanged.
- Equal initial preferences produce equal initial action-selection probabilities, although later preference changes can make the probabilities unequal.