Concepts / Soft-max Action Selection

Soft-max Action Selection

Numerical preferences express how strongly each action is favored.

  • Programming

From Favoring to Choosing

An action-selection method needs a way to express which actions are currently favored. In gradient bandit methods, that role is handled by a numerical preference for every action. The preference indicates how strongly an action is favored, but it is not an estimate of the reward that action will produce.

A preference is a relative favoring signal. The preference itself is not the final decision and is not an estimate of reward.

The Three-Stage Selection Process

Soft-max action selection can be understood as a three-stage process. First, each action a has a numerical preference H_t(a) at time t. Second, the preferences are processed by a soft-max distribution. Third, the resulting probability π_t(a) determines how likely that action is to be taken at time t. A larger preference makes an action more strongly favored relative to actions with smaller preferences.

processproducedetermine likelihoodH_t(a)action preferencesSoft-maxpreference distributionπ_t(a)action probabilitiesActionselected according toprobability
How does a set of numerical preferences become action probabilities that determine which action is selected?

The soft-max distribution is therefore the link between an internal preference state and an action choice. It does not simply select the action with the largest preference as a fixed rule. Instead, it maps the collection of preferences to probabilities, so the relative preferences determine how strongly each action is favored.

Equal Preferences at the Start

Three actions begin equally favored

Suppose three actions begin with equal numerical preferences.

Record the preferences: Each action has the same initial preference, so no action is relatively more favored than another.

Apply soft-max: The soft-max distribution receives equal preferences and therefore assigns equal initial probabilities to the actions.

Interpret the result: The actions begin equally likely to be selected. This is a starting condition, not a guarantee that their probabilities will remain equal.

Equal initial preferences produce equal initial action-selection probabilities.

soft-maxsoft-maxsoft-maxAction Aequal preferenceπ_t(A)equal probabilityAction Bequal preferenceπ_t(B)equal probabilityAction Cequal preferenceπ_t(C)equal probability
When all actions begin with the same preference, how does soft-max distribute their initial selection probabilities?

Why Absolute Values Do Not Matter

Soft-max action selection depends on relative preferences rather than on the absolute numerical labels attached to them. If one action has a larger preference than another, it is more strongly favored relative to that other action. What matters is this relationship between preferences.

soft-maxsoft-maxsoft-maxsoft-maxH_t(A)Action Aπ_t(a)original probabilitiesH_t(A) + 1000Action Aπ_t(a)same probabilitiesH_t(B)Action BH_t(B) + 1000Action B
What happens to action probabilities when the same constant is added to every preference?

For example, adding 1000 to every preference changes the numerical labels but does not change the relative preferences. Because soft-max selection depends on those relative preferences, the action probabilities remain unchanged. This is why a preference value should be interpreted in comparison with the other preferences, not in isolation.

Mistakes About Preference Values

  • Treating H_t(a) as an estimate of reward

    In gradient bandit methods, the preference indicates how strongly an action is favored, but it is not itself an estimate of reward.

    Fix: Interpret H_t(a) as a preference signal that soft-max converts into an action probability.

  • Thinking the largest preference is the entire selection rule

    The preferences are processed by a soft-max distribution, which produces probabilities for taking actions.

    Fix: Follow the full process: preferences are transformed into probabilities, and those probabilities determine how likely each action is to be taken.

  • Believing that equal initial probabilities remain equal

    The initial equality only reflects equal initial preferences. Later relative changes can produce new action probabilities.

    Fix: Reconsider the relative preferences whenever the preference values change.

  • Assuming that adding the same constant changes the probabilities

    The common shift changes absolute values but leaves relative preferences unchanged.

    Fix: Compare the preferences with one another; a common constant does not affect the resulting action probabilities.

Check Your Understanding

EASY

Imagine that every action in a selection problem receives the same increase in preference. Will the soft-max action probabilities change? Explain your answer using the idea of relative preferences.

Hints
  • Compare each action's preference with the other actions' preferences before and after the increase.
  • Ask whether the ordering or differences in relative favoring have changed.

What do you think happens?

Three actions begin with equal preferences. Before reading the reveal, predict whether their initial soft-max probabilities are equal or unequal.

  • Equal
  • Unequal
Reveal answer

Answer: Equal

Equal initial preferences mean that no action is relatively more favored than another, so soft-max produces equal initial action-selection probabilities.

Key Takeaways

  • H_t(a) is the numerical preference for action a at time t.
  • A preference expresses how strongly an action is favored; it is not itself an estimate of reward.
  • Soft-max transforms the collection of preferences into action probabilities π_t(a).
  • Relative preferences determine the probabilities, so adding the same constant to every preference leaves them unchanged.
  • Equal initial preferences produce equal initial action-selection probabilities, although later preference changes can make the probabilities unequal.