Action-Value Estimation
UCB combines an estimated action value with an uncertainty bonus.
Why Estimated Value Is Not Enough
Suppose several actions are available and each has an estimated value. Choosing only the action with the highest current estimate can overlook an action whose estimate is uncertain. Upper Confidence Bound action selection addresses this by considering an upper confidence bound for each action's estimated value. It favors the action with the largest bound, not necessarily the action with the largest estimate alone.
Upper Confidence Bound action selection combines an estimated action value with an uncertainty bonus, then selects the action whose combined quantity is greatest.
The Two Parts of a Bound
The UCB quantity has two main parts. The estimated value represents what is currently believed about action a. The added square-root term represents uncertainty, or variance, in that estimate. The action with the greatest combined quantity is favored.
UCB for action a = estimated value of action a + c × square-root uncertainty term for action a
The square-root term is not another estimate of the action's value. It is an estimate of how uncertain the value estimate is. A larger uncertainty contribution can raise an action's upper confidence bound even when its current estimated value is not the largest.
Following the Selection Counts
The uncertainty term changes after each selection. Let N_t(a) denote the number of times action a has been selected by time t. When action a itself is selected, N_t(a) is incremented. Since N_t(a) appears in the denominator of the square-root term, the uncertainty estimate for action a decreases.
When a different action is selected, time t increases but N_t(a) does not. The numerator therefore increases while action a's denominator remains unchanged, so the uncertainty estimate for action a increases. In this way, an action that is skipped can become more attractive through its growing uncertainty bonus.
What do you think happens?
Action a is selected once, while action b is not selected during the same step. What happens to their uncertainty estimates?
Reveal answer
Answer: Action a decreases and action b increases.
Selecting action a increments N_t(a), increasing the denominator of its square-root term. For action b, time t increases while N_t(b) remains unchanged, so its uncertainty estimate increases.
Worked Selection Trace
Two Actions Across Two Steps
Consider two actions, a and b. At the beginning, action a has been selected more often than action b. First select action a, then select action a again. Track the direction of change in each action's uncertainty estimate.
Before the first selection: Action b has the smaller previous selection count in this generated scenario, so its uncertainty estimate is the larger one. Its upper confidence bound may therefore receive a stronger uncertainty contribution.
After selecting action a once: The selection count N_t(a) increases. Because that count is in the denominator of action a's square-root term, action a's uncertainty estimate decreases. Action b is skipped: time increases, its selection count does not, and its uncertainty estimate increases.
After selecting action a again: Action a's selection count increases again, so its uncertainty estimate decreases again. Action b is skipped again, so its uncertainty estimate increases again.
Selection implication: Even if action a currently has a higher estimated value, action b's growing uncertainty bonus can raise action b's upper confidence bound. UCB compares the combined bounds rather than estimated values alone.
Selecting an action reduces that action's uncertainty estimate, while repeatedly skipping another action increases the skipped action's uncertainty estimate.
| Event | Selected action's count | Skipped action's count | Selected action's uncertainty | Skipped action's uncertainty |
|---|---|---|---|---|
| Select action a | Increases | Unchanged | Decreases | Increases |
| Select action a again | Increases again | Unchanged again | Decreases again | Increases again |
Qualitative changes in the uncertainty estimates after repeatedly selecting action a.
The Role of c
The parameter c is greater than zero. It controls the degree of exploration and determines the confidence level of the upper bound. Because c multiplies the uncertainty term, changing c changes how strongly uncertainty affects the quantity being maximized.
| Parameter choice | Effect on uncertainty bonus | Effect on selection pressure |
|---|---|---|
| Larger c | The uncertainty term has a stronger effect | More emphasis on exploration and a higher confidence level for the upper bound |
| Smaller positive c | The uncertainty term has a weaker effect | Less emphasis on exploration and a lower confidence level for the upper bound |
Untried Actions
If N_t(a) equals zero, action a has not been selected before time t. The action has no previous selection count available for the denominator of the uncertainty term. UCB handles this case by treating the action as a maximizing action, which makes the untried action eligible for selection rather than allowing the missing count to block the comparison.
Mistakes in UCB Reasoning
Selecting only the action with the highest estimated value
An action with a lower current estimate may have greater uncertainty and therefore a larger upper confidence bound.
Fix:
Compare the estimated value plus the uncertainty bonus.Assuming a skipped action becomes less uncertain
When another action is selected, time t increases while the skipped action's selection count remains unchanged. Its uncertainty estimate increases.
Fix:
Track both time and the action-specific selection count.Treating c as part of the estimated action value
c multiplies the uncertainty term and changes how strongly uncertainty affects the quantity being maximized.
Fix:
Keep the estimated value and the scaled uncertainty bonus conceptually separate.Ignoring an action whose selection count is zero
UCB treats an action with N_t(a) equal to zero as a maximizing action.
Fix:
Handle the zero-count case explicitly so the untried action remains eligible for selection.
Check Your Understanding
Two actions have the same estimated value. Action a has been selected many times, while action b has been selected only a few times. Which action receives the larger uncertainty contribution, and what additional factor determines how strongly that contribution affects the UCB quantity?
Hints
- Look at the selection count in the denominator of the square-root term.
- The parameter c multiplies the uncertainty term.
Action b is skipped while action a is selected. Predict the direction of change in action b's uncertainty estimate and explain the roles of t and N_t(b).
Hints
- Time t increases.
- N_t(b) remains unchanged because b was not selected.
Key Takeaways
- UCB combines an estimated action value with an uncertainty bonus and favors the largest combined bound.
- The square-root term estimates uncertainty or variance in an action's value estimate.
- Selecting an action increases its selection count and decreases its uncertainty estimate.
- Skipping an action leaves its selection count unchanged while time increases, so its uncertainty estimate increases.
- The positive parameter c controls exploration and the confidence level by scaling the uncertainty term.
- An action with zero previous selections is treated as a maximizing action so it remains eligible for selection.
Key Takeaways
- Upper Confidence Bound selection weighs both what is currently estimated and what remains uncertain.
- The square-root term represents uncertainty, while c controls how strongly that uncertainty affects selection.
- Selecting an action reduces its uncertainty estimate; selecting other actions increases the uncertainty estimate of an action that was skipped.
- An action with no previous selections is treated as a maximizing action so that it can be selected.