Greedy and Nongreedy Actions
Exploitation uses the action that currently has the greatest estimated value.
The Choice Behind Every Action
At each time step, an action-selection method must decide what to do next. It can select the action that currently appears best, or it can select another action to learn more about that action's value. These choices are called exploitation and exploration. The tension between them is the exploration-exploitation trade-off.
Greedy and Nongreedy Choices
A greedy action is the action that currently has the greatest estimated value. Choosing it is exploitation: the agent uses what it currently believes and expects to receive the greatest reward on this step. A nongreedy action is an action that is not currently estimated to be the best. Choosing it is exploration when the purpose is to improve knowledge about that action's value.
Reading the current estimates
Suppose an agent currently estimates that Action A has a greater value than Action B.
Identify the greedy action: Action A is greedy because it currently has the greatest estimated value.
Identify exploitation: Choosing Action A exploits current knowledge and aims for the greatest expected reward on this step.
Identify exploration: Choosing Action B is nongreedy. If the purpose is to improve knowledge about Action B's value, that choice is exploration.
The same action-selection situation supports two different purposes: choose Action A to exploit, or choose Action B to explore.
Immediate and Long-Run Rewards
Exploitation is appropriate when the goal is to maximize expected reward on the one current step. However, the action that currently looks best may not truly be the best action. Exploration accepts a lower reward in the short run when it may improve knowledge about another action. If that investigation reveals a better action, the newly identified action can be exploited repeatedly, producing greater total reward over many time steps.
Comparing two objectives
An agent believes Action A is best right now, but Action B has not been investigated enough to know whether it might be better.
One-step objective: If the agent only wants the greatest expected reward on the current step, it chooses the currently best-looking action, Action A.
Long-run objective: If the agent values total reward over many steps, it may choose Action B to improve its estimate, even if this creates a lower reward in the short run.
Possible later benefit: If investigating Action B reveals that it is better than Action A, the agent can exploit Action B repeatedly on later steps.
The best choice for one immediate reward and the best choice for total reward over many steps can differ.
Uncertainty Changes the Trade-Off
An estimated value is not a guarantee that an action is truly best. The currently highest estimate may be wrong because the agent's knowledge is incomplete. This possibility makes a nongreedy action worth considering: trying it can improve the estimate and reveal information that is not available from repeatedly choosing the current favorite.
Common Reasoning Mistakes
Treating the greedy action as certainly the best action
The action that currently looks best may not truly be the best action.
Fix:
Remember that exploration can improve knowledge about a nongreedy action and may reveal a better choice.Calling every nongreedy choice exploration
Exploration specifically involves choosing a nongreedy action to improve knowledge about its value.
Fix:
Use the term exploration when the nongreedy choice is intended to investigate or improve the value estimate.Assuming the action with the best immediate reward must produce the best total reward
Immediate reward and long-run total reward can favor different choices.
Fix:
Separate the one-step objective from the possibility that exploration can lead to greater reward across many later steps.Claiming that one action can explore and exploit at the same time
A single action selection cannot both choose the current best option and choose another option to improve its estimate.
Fix:
At each decision point, identify whether the selected action is being used for exploitation or for exploration.
Practice the Decision
An agent currently estimates that Action A has the greatest value. It chooses Action B because it wants to learn whether Action B may be better over many future steps. Identify the action-selection purpose, explain whether the chosen action is greedy or nongreedy, and state why the choice may still improve total reward.
Hints
- Compare the chosen action with the action that currently has the greatest estimated value.
- Ask whether the choice is using current knowledge or improving knowledge.
- Distinguish the reward on this step from reward accumulated over many steps.
What do you think happens?
If Action A currently has the greatest estimated value and the agent chooses Action B specifically to improve knowledge about B, is the decision exploitation or exploration?
Reveal answer
Answer: Exploration
Action B is nongreedy because it is not currently estimated to be the best, and choosing it to improve knowledge about its value is exploration. One action selection cannot simultaneously explore and exploit.
Key Takeaways
- Exploitation chooses the greedy action: the action with the greatest current estimated value.
- Exploration chooses a nongreedy action to improve knowledge about its value.
- A single action selection cannot simultaneously exploit the current best-looking action and explore another action.
- Exploitation can maximize expected reward on the current step, while exploration may support greater total reward over many steps.
- Uncertainty matters because the current favorite may not truly be the best action.
Key Takeaways
- A greedy action has the greatest current estimated value, so choosing it is exploitation.
- A nongreedy action can be chosen for exploration when the goal is to improve knowledge about its value.
- Immediate expected reward and long-run total reward can favor different actions.
- Exploration can be worthwhile when current estimates are uncertain and learning may reveal a better action.
- At one decision point, the agent must choose between exploiting current knowledge and exploring another action.