Action-Value Methods for Action Selection
Numerical preferences express how strongly each action is favored.
From Preference to Action
An action-selection method must decide how likely each available action is to be chosen. In this approach, the method does not begin by treating its internal numbers as direct reward estimates. Instead, it learns a numerical preference for every action. The preference records how strongly the action is favored relative to the other available actions.
A preference is a measure of favor, not itself an estimate of reward.
The Three-Stage Selection Process
The method can be understood as a three-stage process. First, each action a has a numerical preference H_t(a) at time t. Second, the collection of preferences is processed by a soft-max distribution. Third, the soft-max result gives π_t(a), the probability that action a will be selected at time t. A larger preference makes an action more strongly favored relative to actions with smaller preferences.
Tracing Three Actions
Suppose three available actions have preferences H_t(A) = 2, H_t(B) = 5, and H_t(C) = 1. Which action is favored most before the probabilities are considered?
Identify the preferences: The method associates one numerical preference with each action: A has 2, B has 5, and C has 1.
Compare the preferences: Action B has the largest preference, so it is favored more strongly than A and C. Action C has the smallest preference.
Apply soft-max conceptually: The soft-max distribution uses these relative preferences to produce an action-selection probability for each action. The action with the larger preference is more strongly favored relative to the others.
The preference ordering is B, then A, then C. The resulting probability ordering follows that preference ordering: B is most favored, while C is least favored.
Relative Values Control Choice
The absolute size of a preference is less important than its position relative to the other preferences. If every preference is increased by the same constant, the numerical labels change, but the relative preferences do not. Because soft-max selection depends on those relative preferences, the action probabilities remain unchanged.
Adding the Same Constant
Compare preferences H_t(A) = 1, H_t(B) = 3, and H_t(C) = 2 with preferences obtained by adding 1000 to every value.
Record the original ordering: The original ordering is B, then C, then A because 3 is greater than 2, which is greater than 1.
Add the same constant: The new preferences are 1001, 1003, and 1002. The numerical values are much larger, but their ordering is still B, then C, then A.
Compare the selection probabilities: The soft-max probabilities remain unchanged because adding the same constant does not alter the relative preferences.
The preference labels change, but the action-selection probabilities do not.
Equal Starting Preferences
When all actions begin with equal preferences, no action is favored over another by the preference values. The soft-max distribution therefore produces equal initial probabilities for the actions. This is only a starting condition: later changes to the numerical preferences can make the actions no longer equally likely.
What do you think happens?
Three actions start with equal preferences. Before any preferences change, what should you expect about their initial selection probabilities?
Reveal answer
Answer: The actions have equal initial probabilities.
Equal initial preferences do not favor one action over another, so soft-max produces equal initial action-selection probabilities.
Changing Preferences Over Time
Equal starting probabilities do not imply that the actions remain equally likely. As the numerical preferences change, their relative ordering can change as well. The soft-max distribution then uses the new relative preferences to determine new action probabilities. The important state to trace is therefore not only the current value for one action, but the relationship among all action preferences.
Tracing a Preference Update
Three actions begin with equal preferences. Later, action B has a larger preference than actions A and C. What changes?
Initial state: Because the initial preferences are equal, the initial action-selection probabilities are equal.
Preference change: Action B becomes more strongly favored relative to A and C when its preference becomes larger than theirs.
Probability change: The soft-max distribution processes the new relative preferences, so B becomes more strongly favored in the resulting action probabilities.
Equal probabilities describe the initial state only. A change in relative preferences produces a new probability distribution.
When tracing this method, write the actions and their preferences as pairs, compare the preferences before applying soft-max conceptually, and then describe how the probability ordering follows the relative preference ordering.
Common Interpretation Errors
Treating H_t(a) as a direct estimate of reward.
The preference indicates how strongly the action is favored, but it is not itself an estimate of reward.
Fix:
Interpret H_t(a) as a numerical preference that is later converted into action-selection probabilities.Comparing absolute preference values without considering the other actions.
Adding the same constant to every preference changes the labels but not the relative preferences or action probabilities.
Fix:
Compare each preference with the other preferences in the same selection state.Assuming equal initial probabilities continue forever.
The equal-probability result applies to equal initial preferences. Later preference changes can alter the probability distribution.
Fix:
Reevaluate the relative preferences whenever the numerical preferences change.Skipping the soft-max stage.
Preferences are processed by soft-max before they determine how likely each action is to be taken.
Fix:
Trace the full sequence: H_t(a), then soft-max, then π_t(a).
Check Your Understanding
Two actions have preferences H_t(A) = 6 and H_t(B) = 2. Explain which action is more strongly favored, what role soft-max plays, and whether adding 50 to both preferences changes the action probabilities.
Hints
- Compare the two preferences rather than either value in isolation.
- Separate the preference stage from the probability stage.
- Adding the same constant preserves the relative preferences.
- The method gives every action a numerical preference H_t(a). Soft-max processes those preferences and produces π_t(a), the probability of selecting each action at time t. Larger preferences favor actions relative to smaller preferences. Equal initial preferences produce equal initial probabilities, but later changes in relative preferences can change the probabilities. Adding the same constant to every preference changes the numerical values without changing the action probabilities.
Key Takeaways
- H_t(a) expresses how strongly action a is favored at time t.
- The preference is not itself an estimate of reward.
- Soft-max converts the collection of preferences into action-selection probabilities π_t(a).
- Relative preferences determine the probabilities, so adding the same constant to every preference leaves them unchanged.
- Equal initial preferences produce equal initial probabilities, but later preference changes can make the actions unequally likely.