Concepts / Balancing Exploration and Exploitation with Upper-Confidence-Bound Action Selection

Balancing Exploration and Exploitation with Upper-Confidence-Bound Action Selection

UCB combines an estimated action value with an uncertainty bonus.

  • Programming

The Choice Between Discovery and Reward

A reinforcement learning agent often must choose an action before it knows which action is best. It can exploit an action that has already produced reward, or it can explore an unfamiliar action that might reveal a better opportunity. These competing needs create the exploration-exploitation dilemma.

Exploration is trying unfamiliar actions to discover how effective they are. Exploitation is using actions that have already been found to be effective.

Neither behavior is sufficient by itself. Exclusive exploration keeps spending decisions on discovery instead of using actions that appear effective. Exclusive exploitation repeatedly uses the current favorite and may never check whether another action is better. An agent therefore needs discovery to gain information and reward-seeking behavior to make use of that information.

supplies experienceuses effective actionsExplorationdiscover effectivenessLearningdiscovery andreward-seekingExploitationuse effective actions
How does the agent balance trying unfamiliar actions with using actions that already appear effective?

How UCB Builds an Action Score

Upper Confidence Bound action selection addresses the weakness of choosing only the action with the highest current estimated value. A high estimate may belong to an action whose value is still uncertain. UCB considers an upper confidence bound for each action and favors the action with the largest combined quantity.

A UCB score combines two parts: the estimated value of an action and an added uncertainty bonus. The estimated value represents what is currently believed about the action. The uncertainty term represents uncertainty, or variance, in that estimate.

contributescontributesEstimated valuecurrent belief+addedSquare-root termuncertainty bonusUCB scorecombined quantity
How do the estimated reward and square-root uncertainty term combine to produce an action's UCB score?

The square-root term is important because it prevents the agent from treating an uncertain estimate as if it were fully reliable. An action can therefore be selected because its estimated value is high, because its uncertainty bonus is high, or because the combination is highest among the available actions.

Tracing Selection Counts Through Time

The uncertainty term changes as selections are made. For action a, the count N_t(a) records how often that action has been selected by time t. When action a is selected, N_t(a) increases. Because this count is in the denominator of the square-root term, the uncertainty estimate for that action decreases.

When a different action is selected, time t increases but N_t(a) does not. The numerator of the uncertainty term increases while action a's denominator remains unchanged, so action a's uncertainty estimate increases. In this way, an action that is skipped becomes more attractive to reconsider, while an action that is repeatedly selected becomes less uncertain.

selectedreducesnot selectedincreasesAction acurrent count N_t(a)Action a skippedN_t(a) unchangedAction a selectedN_t(a) increasesHigher uncertaintytime increasesLower uncertaintydenominator increases
What happens to an action's uncertainty estimate when it is selected again, while the uncertainty of skipped actions remains larger?

Selected Versus Skipped

Compare what happens to two actions after the agent selects action A and skips action B.

Action A is selected: The selection count for A increases. Since its count is in the denominator of the uncertainty term, A's uncertainty estimate decreases.

Action B is skipped: The overall time increases, but B's selection count stays the same. The numerator grows while B's denominator does not, so B's uncertainty estimate increases.

Compare the estimates: A now has more direct selection experience, while B has become relatively more uncertain and may receive more exploration pressure.

Selecting an action reduces that action's uncertainty estimate; skipping it while other actions are selected increases its uncertainty estimate.

The Role of the Exploration Parameter

The parameter c is greater than zero. It multiplies the uncertainty term, so it controls how strongly uncertainty affects the quantity being maximized. It also controls the degree of exploration and determines the confidence level of the upper bound.

tends towardtends towardSmaller cuncertainty has lessinfluenceLess explorationrelative effectLarger cuncertainty has moreinfluenceMore explorationrelative effect
How does increasing or decreasing c change the size of the uncertainty bonus and the agent's willingness to explore?

Comparing Two Settings of c

Suppose two actions have the same estimated value, but one action is more uncertain. What changes when c is made larger?

Start with equal estimated values: The estimated-value parts do not distinguish the actions.

Increase c: Because c multiplies the uncertainty term, the uncertain action's bonus has a stronger effect on the combined quantity.

Interpret the result: The action with greater uncertainty receives more influence from the bonus, so the agent becomes more willing to explore it.

Increasing c gives uncertainty more weight and corresponds to a stronger exploration preference and a different confidence level for the upper bound.

Why Untried Actions Are Eligible

If N_t(a) equals zero, action a has not been selected before time t. The uncertainty expression would not have a usable previous selection count in its denominator. UCB handles this case by considering the action a maximizing action, which ensures that an untried action is eligible for selection.

has selection historyspecial handlingTried actionN_t(a) greater than zeroAvailable estimateprevious selection countexistsUntried actionN_t(a) equals zeroMaximizing actioneligible for selection
Why does an action with zero selections receive an effectively infinite uncertainty bonus and therefore get selected?

An Action with No History

An agent has several actions with selection histories and one action whose count is zero. How should UCB treat the action with zero selections?

Inspect the count: The action's count N_t(a) is zero, so the agent has no previous selection count for that action.

Apply the edge-case rule: UCB treats the action as a maximizing action rather than relying on an unavailable denominator.

Interpret the choice: The untried action is made eligible for selection, allowing the agent to discover its effectiveness.

An action with no previous selections is treated as a maximizing action so that it can be selected and evaluated.

Repeated Trials in Uncertain Tasks

A single observed reward does not provide dependable information about an action when the task is stochastic. Repeated selections supply experience about that action. As the agent gathers more information, its estimate can become more reliable, while the uncertainty term reflects how much uncertainty remains.

gather more observationssupportsFirst selectionlimited experienceRepeated selectionsmore experienceAction-value estimatemore reliable information
How do repeated selections of the same action cause noisy observed rewards to produce a more reliable estimated value?

This is why an agent's preference should develop progressively. Early exploration supplies experience. Later choices can give greater weight to actions that appear effective, while continued exploration remains important whenever the agent still lacks reliable information.

Mistakes in UCB Reasoning

  • Choosing only the action with the highest estimated value.

    UCB is designed to consider the uncertainty bonus as well as the estimated value.

    Fix: Compare the combined upper confidence quantities rather than the estimated values alone.

  • Assuming uncertainty always decreases over time.

    When another action is selected, time increases while the skipped action's selection count stays unchanged, so its uncertainty estimate increases.

    Fix: Track both time and the action's own selection count.

  • Treating c as an action value.

    c controls the influence of the uncertainty term and the confidence level of the upper bound.

    Fix: Keep the estimated value and the exploration-controlling parameter conceptually separate.

  • Ignoring an action whose selection count is zero.

    UCB treats an action with zero previous selections as a maximizing action so that it remains eligible for selection.

    Fix: Handle the zero-count case explicitly.

  • Using only exploration or only exploitation.

    Exclusive exploration fails to use effective actions, while exclusive exploitation can miss a better action.

    Fix: Allow exploration to supply information and exploitation to use actions that appear effective.

Practice: Predict the Next Preference

MEDIUM

An agent has selected action A several times. Action B has been skipped while other actions were selected, so B's selection count has not increased. Without calculating a numerical score, predict how the uncertainty estimates for A and B change relative to their previous values. Then explain why UCB might give B more consideration even if A currently has the higher estimated value.

Hints
  • Ask what happens to an action's denominator when that action is selected.
  • For the skipped action, ask what happens to time while its own selection count stays unchanged.
  • Remember that UCB compares estimated value together with uncertainty.

What do you think happens?

What happens to an action's uncertainty estimate when another action is selected and its own selection count does not change?

  • It decreases
  • It increases
  • It stays unchanged
Reveal answer

Answer: It increases.

Time increases while the action's selection count remains unchanged. The numerator therefore increases while its denominator stays the same.

What to Remember

  1. UCB combines an estimated action value with an uncertainty bonus and selects the action with the largest combined upper confidence quantity.
  2. The square-root term represents uncertainty or variance in the action-value estimate.
  3. Selecting an action increases its selection count and decreases its uncertainty estimate; selecting other actions increases the uncertainty estimate of an action that was skipped.
  4. The parameter c is greater than zero and controls how strongly uncertainty affects the score, along with the exploration degree and confidence level.
  5. An action with zero previous selections is treated as a maximizing action so that it remains eligible for discovery.
  6. Exploration discovers action effectiveness, while exploitation uses actions already found to be effective; reinforcement learning needs both.

Key Takeaways

  • UCB balances exploitation of estimated rewards with exploration of uncertain actions.
  • Its uncertainty bonus changes with selection history: selected actions become less uncertain, while skipped actions become more uncertain.
  • The parameter c controls the influence of uncertainty and the confidence level of the upper bound.
  • Untried actions are treated as maximizing actions so that the agent can discover their effectiveness.
  • Repeated trials provide experience needed to make action-value estimates more reliable in uncertain tasks.