Concepts / Exploration in Action Selection

Exploration in Action Selection

Action-value estimates are uncertain, so the action that currently looks best may not actually be best.

  • Programming

When the Best-Looking Action Is Wrong

When an agent chooses an action, it usually has an estimate of how valuable each available action is. These estimates are uncertain. As a result, the action with the highest current estimate may not actually be the best action. Exploration addresses this problem by giving the agent a reason to investigate actions whose values may have been underestimated.

estimate may be wrongestimate may be wrongAction Ahighest estimateAction Alower valueAction BunderestimatedAction Bhigher value
How can the action that currently has the highest estimate turn out not to be the truly best action?

Greedy and Non-Greedy Choices

The greedy action is the action with the highest current estimate. Selecting it uses the information the agent currently has. A non-greedy action is any available action that does not have the highest current estimate. Selecting one can provide information about whether its estimate is wrong and whether its actual value may be higher than it currently appears.

ChoiceWhat it followsWhy it may be selected
Greedy actionThe highest current estimateIt currently looks best
Non-greedy actionAn action below the highest current estimateIts estimate may be wrong or underestimated

The important distinction is between using an estimate and investigating an estimate. A greedy choice asks, “Which action looks best now?” An exploratory choice asks, “Could another action actually be better despite its lower current estimate?” Upper-confidence-bound selection is based on the second question.

adds uncertaintymay become attractiveGreedy selectionhighest estimateNon-greedy actionmay be underestimatedUCB selectionestimate plus uncertainty
What changes when action selection considers uncertainty in addition to the current estimate?

Tracing an Uncertain Decision

Two actions with incomplete information

An agent has two available actions. Action A currently has the higher estimated value, while Action B has a lower estimate that may be underestimated.

Read the current estimates: Action A is the greedy action because it has the highest current estimate. Action B is non-greedy because its current estimate is lower.

Notice the uncertainty: The estimates are not guaranteed to be correct. Action B may have a value that is higher than its current estimate suggests.

Compare the decision questions: A purely greedy decision selects Action A because it looks best now. An exploratory decision also considers whether investigating Action B could reveal that it is actually better.

Choose an action-selection principle: A method that accounts for both estimated value and uncertainty can give Action B a reason to be selected, even though Action A currently has the higher estimate.

The current leader is not automatically the truly best action. Exploration is needed because a non-greedy action may turn out to be better.

Upper-Confidence-Bound Selection

Upper-confidence-bound action selection considers both an action's estimated value and its uncertainty. An action can therefore be attractive for two different reasons: it may currently appear valuable, or its uncertain estimate may leave open the possibility that it is better than it appears. This combines the exploitation of promising actions with exploration of actions whose estimates may be unreliable.

contributescontributesguides selectionEstimated valuewhat looks good nowUCB assessmentvalue plus uncertaintySelected actionhighest combined assessmentUncertaintywhat may be underestimated
How does upper-confidence-bound selection combine an action's estimated value with its uncertainty?

The word upper refers to evaluating an action using an optimistic view of what its value could be when uncertainty is taken into account. The method does not simply ask which estimate is largest. It asks which action has the strongest assessment after estimated value and uncertainty are considered together.

The ε-Greedy Trade-Off

ε-greedy selection divides action choices into two broad cases. It usually selects the greedy action, the one with the highest current estimate. For the exploratory case, it selects a non-greedy action indiscriminately. This creates a simple balance between using the current best estimate and trying alternatives.

usual choiceexploration choiceAction selectionGreedy actionhighest estimateNon-greedy actionselected indiscriminately
How does ε-greedy selection split decisions between the best-known action and randomly selected alternatives?
MethodHow it evaluates the current leaderHow it treats non-greedy actions
Greedy selectionChooses the action with the highest current estimateDoes not select them
ε-greedy selectionUsually chooses the action with the highest current estimateTries them indiscriminately
Upper-confidence-bound selectionConsiders the estimated valueAlso considers uncertainty

Why Random Alternatives Are Limited

ε-greedy exploration has an important limitation: when it chooses to explore, it treats non-greedy actions indiscriminately. It does not use the different estimated values of those alternatives to decide which non-greedy action is more promising. Consequently, two non-greedy actions can receive the same exploratory treatment even when one has a much more promising estimate than the other.

  • Assuming that the greedy action is guaranteed to be the truly best action

    Action-value estimates are uncertain, so another action may be underestimated and may actually be better

    Fix: Recognize that exploration is needed to investigate actions whose estimates may be wrong

  • Treating every non-greedy action as equally informative

    ε-greedy selection tries non-greedy actions indiscriminately

    Fix: Separate ε-greedy exploration from uncertainty-aware selection such as upper-confidence-bound selection

  • Thinking that exploration ignores estimated value completely

    Upper-confidence-bound selection considers both estimated value and uncertainty

    Fix: Explain that uncertainty is added to the estimated-value perspective rather than replacing it

When comparing action-selection methods, ask two separate questions: how does the method use the highest current estimate, and how does it choose among actions that are not currently greedy? This exposes the difference between greedy selection, ε-greedy exploration, and uncertainty-aware selection.

Decision Practice

MEDIUM

An agent has three actions. Action A has the highest current estimate. Action B has a lower estimate but substantial uncertainty. Action C also has a lower estimate, but its estimate is less uncertain. Explain which action greedy selection would choose, what ε-greedy exploration does not distinguish between, and why upper-confidence-bound selection may give Action B special consideration.

Hints
  • Start by identifying the action with the highest current estimate.
  • Recall that ε-greedy selection treats non-greedy actions indiscriminately during exploration.
  • For upper-confidence-bound selection, consider both estimated value and uncertainty.
  1. A strong answer identifies Action A as the greedy choice, explains that ε-greedy exploration does not distinguish between Action B and Action C as non-greedy alternatives, and explains that upper-confidence-bound selection can favor Action B because its uncertainty may indicate that its value is underestimated.

Key Takeaways

  • Action-value estimates are uncertain, so the action with the highest current estimate may not be truly best.
  • The greedy action has the highest current estimate; non-greedy actions may be worth investigating because their values may be underestimated.
  • Upper-confidence-bound selection considers both estimated value and uncertainty.
  • ε-greedy selection explores non-greedy actions indiscriminately, so it does not distinguish among their different estimates.