Exploration in Action Selection
Action-value estimates are uncertain, so the action that currently looks best may not actually be best.
When the Best-Looking Action Is Wrong
When an agent chooses an action, it usually has an estimate of how valuable each available action is. These estimates are uncertain. As a result, the action with the highest current estimate may not actually be the best action. Exploration addresses this problem by giving the agent a reason to investigate actions whose values may have been underestimated.
Greedy and Non-Greedy Choices
The greedy action is the action with the highest current estimate. Selecting it uses the information the agent currently has. A non-greedy action is any available action that does not have the highest current estimate. Selecting one can provide information about whether its estimate is wrong and whether its actual value may be higher than it currently appears.
| Choice | What it follows | Why it may be selected |
|---|---|---|
| Greedy action | The highest current estimate | It currently looks best |
| Non-greedy action | An action below the highest current estimate | Its estimate may be wrong or underestimated |
The important distinction is between using an estimate and investigating an estimate. A greedy choice asks, “Which action looks best now?” An exploratory choice asks, “Could another action actually be better despite its lower current estimate?” Upper-confidence-bound selection is based on the second question.
Tracing an Uncertain Decision
Two actions with incomplete information
An agent has two available actions. Action A currently has the higher estimated value, while Action B has a lower estimate that may be underestimated.
Read the current estimates: Action A is the greedy action because it has the highest current estimate. Action B is non-greedy because its current estimate is lower.
Notice the uncertainty: The estimates are not guaranteed to be correct. Action B may have a value that is higher than its current estimate suggests.
Compare the decision questions: A purely greedy decision selects Action A because it looks best now. An exploratory decision also considers whether investigating Action B could reveal that it is actually better.
Choose an action-selection principle: A method that accounts for both estimated value and uncertainty can give Action B a reason to be selected, even though Action A currently has the higher estimate.
The current leader is not automatically the truly best action. Exploration is needed because a non-greedy action may turn out to be better.
Upper-Confidence-Bound Selection
Upper-confidence-bound action selection considers both an action's estimated value and its uncertainty. An action can therefore be attractive for two different reasons: it may currently appear valuable, or its uncertain estimate may leave open the possibility that it is better than it appears. This combines the exploitation of promising actions with exploration of actions whose estimates may be unreliable.
The word upper refers to evaluating an action using an optimistic view of what its value could be when uncertainty is taken into account. The method does not simply ask which estimate is largest. It asks which action has the strongest assessment after estimated value and uncertainty are considered together.
The ε-Greedy Trade-Off
ε-greedy selection divides action choices into two broad cases. It usually selects the greedy action, the one with the highest current estimate. For the exploratory case, it selects a non-greedy action indiscriminately. This creates a simple balance between using the current best estimate and trying alternatives.
| Method | How it evaluates the current leader | How it treats non-greedy actions |
|---|---|---|
| Greedy selection | Chooses the action with the highest current estimate | Does not select them |
| ε-greedy selection | Usually chooses the action with the highest current estimate | Tries them indiscriminately |
| Upper-confidence-bound selection | Considers the estimated value | Also considers uncertainty |
Why Random Alternatives Are Limited
ε-greedy exploration has an important limitation: when it chooses to explore, it treats non-greedy actions indiscriminately. It does not use the different estimated values of those alternatives to decide which non-greedy action is more promising. Consequently, two non-greedy actions can receive the same exploratory treatment even when one has a much more promising estimate than the other.
Assuming that the greedy action is guaranteed to be the truly best action
Action-value estimates are uncertain, so another action may be underestimated and may actually be better
Fix:
Recognize that exploration is needed to investigate actions whose estimates may be wrongTreating every non-greedy action as equally informative
ε-greedy selection tries non-greedy actions indiscriminately
Fix:
Separate ε-greedy exploration from uncertainty-aware selection such as upper-confidence-bound selectionThinking that exploration ignores estimated value completely
Upper-confidence-bound selection considers both estimated value and uncertainty
Fix:
Explain that uncertainty is added to the estimated-value perspective rather than replacing it
When comparing action-selection methods, ask two separate questions: how does the method use the highest current estimate, and how does it choose among actions that are not currently greedy? This exposes the difference between greedy selection, ε-greedy exploration, and uncertainty-aware selection.
Decision Practice
An agent has three actions. Action A has the highest current estimate. Action B has a lower estimate but substantial uncertainty. Action C also has a lower estimate, but its estimate is less uncertain. Explain which action greedy selection would choose, what ε-greedy exploration does not distinguish between, and why upper-confidence-bound selection may give Action B special consideration.
Hints
- Start by identifying the action with the highest current estimate.
- Recall that ε-greedy selection treats non-greedy actions indiscriminately during exploration.
- For upper-confidence-bound selection, consider both estimated value and uncertainty.
- A strong answer identifies Action A as the greedy choice, explains that ε-greedy exploration does not distinguish between Action B and Action C as non-greedy alternatives, and explains that upper-confidence-bound selection can favor Action B because its uncertainty may indicate that its value is underestimated.
Key Takeaways
- Action-value estimates are uncertain, so the action with the highest current estimate may not be truly best.
- The greedy action has the highest current estimate; non-greedy actions may be worth investigating because their values may be underestimated.
- Upper-confidence-bound selection considers both estimated value and uncertainty.
- ε-greedy selection explores non-greedy actions indiscriminately, so it does not distinguish among their different estimates.