Concepts / Action-Value Estimation

Action-Value Estimation

UCB combines an estimated action value with an uncertainty bonus.

  • Programming

Why Estimated Value Is Not Enough

Suppose several actions are available and each has an estimated value. Choosing only the action with the highest current estimate can overlook an action whose estimate is uncertain. Upper Confidence Bound action selection addresses this by considering an upper confidence bound for each action's estimated value. It favors the action with the largest bound, not necessarily the action with the largest estimate alone.

Upper Confidence Bound action selection combines an estimated action value with an uncertainty bonus, then selects the action whose combined quantity is greatest.

addaddcompare boundsEstimated valuecurrent beliefUncertainty bonussquare-root termUpper confidenceboundcombined quantitySelected actionlargest bound
How does UCB combine an action's estimated value and uncertainty term to choose which action to select?

The Two Parts of a Bound

The UCB quantity has two main parts. The estimated value represents what is currently believed about action a. The added square-root term represents uncertainty, or variance, in that estimate. The action with the greatest combined quantity is favored.

UCB for action a = estimated value of action a + c × square-root uncertainty term for action a

The square-root term is not another estimate of the action's value. It is an estimate of how uncertain the value estimate is. A larger uncertainty contribution can raise an action's upper confidence bound even when its current estimated value is not the largest.

denominatoraffects numeratormultiply by caddaddEstimated valueaction aN_t(a)selection counttcurrent timeSquare-root termuncertaintyUncertainty bonusc times uncertaintyUCB scoreestimated value plus bonus
How do the estimated value, selection count, total time, and uncertainty bonus map to the final UCB score for each action?

Following the Selection Counts

The uncertainty term changes after each selection. Let N_t(a) denote the number of times action a has been selected by time t. When action a itself is selected, N_t(a) is incremented. Since N_t(a) appears in the denominator of the square-root term, the uncertainty estimate for action a decreases.

When a different action is selected, time t increases but N_t(a) does not. The numerator therefore increases while action a's denominator remains unchanged, so the uncertainty estimate for action a increases. In this way, an action that is skipped can become more attractive through its growing uncertainty bonus.

selectiondenominator increasesother actions are not selectedtime increases, count does notInitial countsaction a has N_t(a)Select action aN_t(a) increasesAction a uncertaintydecreasesSkip action bN_t(b) unchangedAction b uncertaintyincreases
How does the uncertainty estimate for each action change over successive selections when one action is chosen and the others are skipped?

What do you think happens?

Action a is selected once, while action b is not selected during the same step. What happens to their uncertainty estimates?

  • Both decrease
  • Action a decreases and action b increases
  • Action a increases and action b decreases
  • Both remain unchanged
Reveal answer

Answer: Action a decreases and action b increases.

Selecting action a increments N_t(a), increasing the denominator of its square-root term. For action b, time t increases while N_t(b) remains unchanged, so its uncertainty estimate increases.

Worked Selection Trace

Two Actions Across Two Steps

Consider two actions, a and b. At the beginning, action a has been selected more often than action b. First select action a, then select action a again. Track the direction of change in each action's uncertainty estimate.

Before the first selection: Action b has the smaller previous selection count in this generated scenario, so its uncertainty estimate is the larger one. Its upper confidence bound may therefore receive a stronger uncertainty contribution.

After selecting action a once: The selection count N_t(a) increases. Because that count is in the denominator of action a's square-root term, action a's uncertainty estimate decreases. Action b is skipped: time increases, its selection count does not, and its uncertainty estimate increases.

After selecting action a again: Action a's selection count increases again, so its uncertainty estimate decreases again. Action b is skipped again, so its uncertainty estimate increases again.

Selection implication: Even if action a currently has a higher estimated value, action b's growing uncertainty bonus can raise action b's upper confidence bound. UCB compares the combined bounds rather than estimated values alone.

Selecting an action reduces that action's uncertainty estimate, while repeatedly skipping another action increases the skipped action's uncertainty estimate.

EventSelected action's countSkipped action's countSelected action's uncertaintySkipped action's uncertainty
Select action aIncreasesUnchangedDecreasesIncreases
Select action a againIncreases againUnchanged againDecreases againIncreases again

Qualitative changes in the uncertainty estimates after repeatedly selecting action a.

The Role of c

The parameter c is greater than zero. It controls the degree of exploration and determines the confidence level of the upper bound. Because c multiplies the uncertainty term, changing c changes how strongly uncertainty affects the quantity being maximized.

Parameter choiceEffect on uncertainty bonusEffect on selection pressure
Larger cThe uncertainty term has a stronger effectMore emphasis on exploration and a higher confidence level for the upper bound
Smaller positive cThe uncertainty term has a weaker effectLess emphasis on exploration and a lower confidence level for the upper bound
changes emphasischanges emphasisSmaller positive cweaker uncertainty effectLarger cstronger uncertainty effectLess explorationuncertainty matters lessMore explorationuncertainty matters more
How does increasing or decreasing c change the balance between estimated value and the uncertainty bonus?

Untried Actions

If N_t(a) equals zero, action a has not been selected before time t. The action has no previous selection count available for the denominator of the uncertainty term. UCB handles this case by treating the action as a maximizing action, which makes the untried action eligible for selection rather than allowing the missing count to block the comparison.

equalstreat asgreater than zeroevaluateN_t(a)selection count0not selected beforeMaximizing actioneligible for selectionPositive countprevious selections existUpper confidenceboundcompare normally
What happens when UCB evaluates an action whose selection count is zero, and why does it become a maximizing candidate?

Mistakes in UCB Reasoning

  • Selecting only the action with the highest estimated value

    An action with a lower current estimate may have greater uncertainty and therefore a larger upper confidence bound.

    Fix: Compare the estimated value plus the uncertainty bonus.

  • Assuming a skipped action becomes less uncertain

    When another action is selected, time t increases while the skipped action's selection count remains unchanged. Its uncertainty estimate increases.

    Fix: Track both time and the action-specific selection count.

  • Treating c as part of the estimated action value

    c multiplies the uncertainty term and changes how strongly uncertainty affects the quantity being maximized.

    Fix: Keep the estimated value and the scaled uncertainty bonus conceptually separate.

  • Ignoring an action whose selection count is zero

    UCB treats an action with N_t(a) equal to zero as a maximizing action.

    Fix: Handle the zero-count case explicitly so the untried action remains eligible for selection.

Check Your Understanding

MEDIUM

Two actions have the same estimated value. Action a has been selected many times, while action b has been selected only a few times. Which action receives the larger uncertainty contribution, and what additional factor determines how strongly that contribution affects the UCB quantity?

Hints
  • Look at the selection count in the denominator of the square-root term.
  • The parameter c multiplies the uncertainty term.
EASY

Action b is skipped while action a is selected. Predict the direction of change in action b's uncertainty estimate and explain the roles of t and N_t(b).

Hints
  • Time t increases.
  • N_t(b) remains unchanged because b was not selected.

Key Takeaways

  1. UCB combines an estimated action value with an uncertainty bonus and favors the largest combined bound.
  2. The square-root term estimates uncertainty or variance in an action's value estimate.
  3. Selecting an action increases its selection count and decreases its uncertainty estimate.
  4. Skipping an action leaves its selection count unchanged while time increases, so its uncertainty estimate increases.
  5. The positive parameter c controls exploration and the confidence level by scaling the uncertainty term.
  6. An action with zero previous selections is treated as a maximizing action so it remains eligible for selection.

Key Takeaways

  • Upper Confidence Bound selection weighs both what is currently estimated and what remains uncertain.
  • The square-root term represents uncertainty, while c controls how strongly that uncertainty affects selection.
  • Selecting an action reduces its uncertainty estimate; selecting other actions increases the uncertainty estimate of an action that was skipped.
  • An action with no previous selections is treated as a maximizing action so that it can be selected.