Concepts / Greedy and Nongreedy Actions

Greedy and Nongreedy Actions

Exploitation uses the action that currently has the greatest estimated value.

  • Machine Learning

The Choice Behind Every Action

At each time step, an action-selection method must decide what to do next. It can select the action that currently appears best, or it can select another action to learn more about that action's value. These choices are called exploitation and exploration. The tension between them is the exploration-exploitation trade-off.

informchoose current bestchoose anotherAction estimatescurrent valuesAction choiceExploitationgreatest expected rewardnowExplorationimproved knowledge
At a single decision point, how does the selected action commit the agent either to using current knowledge or to gathering information?

Greedy and Nongreedy Choices

A greedy action is the action that currently has the greatest estimated value. Choosing it is exploitation: the agent uses what it currently believes and expects to receive the greatest reward on this step. A nongreedy action is an action that is not currently estimated to be the best. Choosing it is exploration when the purpose is to improve knowledge about that action's value.

exploitexploreGreedy actionhighest current estimateImmediate rewardgreatest expected rewardNongreedy actionnot highest currentestimateAction knowledgeimproved value estimate
How does choosing the action with the highest current estimated value differ from choosing another action to learn more?

Reading the current estimates

Suppose an agent currently estimates that Action A has a greater value than Action B.

Identify the greedy action: Action A is greedy because it currently has the greatest estimated value.

Identify exploitation: Choosing Action A exploits current knowledge and aims for the greatest expected reward on this step.

Identify exploration: Choosing Action B is nongreedy. If the purpose is to improve knowledge about Action B's value, that choice is exploration.

The same action-selection situation supports two different purposes: choose Action A to exploit, or choose Action B to explore.

Immediate and Long-Run Rewards

Exploitation is appropriate when the goal is to maximize expected reward on the one current step. However, the action that currently looks best may not truly be the best action. Exploration accepts a lower reward in the short run when it may improve knowledge about another action. If that investigation reveals a better action, the newly identified action can be exploited repeatedly, producing greater total reward over many time steps.

investigateidentify better choicerepeat exploitationCurrent stepnongreedy actionNew informationvalue estimate improvesLater stepsbetter action exploitedGreater total rewardpossible long-run result
How can an action with lower immediate expected reward lead to greater total reward over many time steps?

Comparing two objectives

An agent believes Action A is best right now, but Action B has not been investigated enough to know whether it might be better.

One-step objective: If the agent only wants the greatest expected reward on the current step, it chooses the currently best-looking action, Action A.

Long-run objective: If the agent values total reward over many steps, it may choose Action B to improve its estimate, even if this creates a lower reward in the short run.

Possible later benefit: If investigating Action B reveals that it is better than Action A, the agent can exploit Action B repeatedly on later steps.

The best choice for one immediate reward and the best choice for total reward over many steps can differ.

Uncertainty Changes the Trade-Off

An estimated value is not a guarantee that an action is truly best. The currently highest estimate may be wrong because the agent's knowledge is incomplete. This possibility makes a nongreedy action worth considering: trying it can improve the estimate and reveal information that is not available from repeatedly choosing the current favorite.

exploitexploreAction Ahighest current estimateAction A knowledgealready favoredAction Blower current estimateAction B knowledgecould improve throughexploration
How can uncertainty in estimated action values make a nongreedy action worth trying even when its current estimate is not the highest?

Common Reasoning Mistakes

  • Treating the greedy action as certainly the best action

    The action that currently looks best may not truly be the best action.

    Fix: Remember that exploration can improve knowledge about a nongreedy action and may reveal a better choice.

  • Calling every nongreedy choice exploration

    Exploration specifically involves choosing a nongreedy action to improve knowledge about its value.

    Fix: Use the term exploration when the nongreedy choice is intended to investigate or improve the value estimate.

  • Assuming the action with the best immediate reward must produce the best total reward

    Immediate reward and long-run total reward can favor different choices.

    Fix: Separate the one-step objective from the possibility that exploration can lead to greater reward across many later steps.

  • Claiming that one action can explore and exploit at the same time

    A single action selection cannot both choose the current best option and choose another option to improve its estimate.

    Fix: At each decision point, identify whether the selected action is being used for exploitation or for exploration.

Practice the Decision

EASY

An agent currently estimates that Action A has the greatest value. It chooses Action B because it wants to learn whether Action B may be better over many future steps. Identify the action-selection purpose, explain whether the chosen action is greedy or nongreedy, and state why the choice may still improve total reward.

Hints
  • Compare the chosen action with the action that currently has the greatest estimated value.
  • Ask whether the choice is using current knowledge or improving knowledge.
  • Distinguish the reward on this step from reward accumulated over many steps.

What do you think happens?

If Action A currently has the greatest estimated value and the agent chooses Action B specifically to improve knowledge about B, is the decision exploitation or exploration?

  • Exploitation
  • Exploration
  • Both at the same time
Reveal answer

Answer: Exploration

Action B is nongreedy because it is not currently estimated to be the best, and choosing it to improve knowledge about its value is exploration. One action selection cannot simultaneously explore and exploit.

Key Takeaways

  1. Exploitation chooses the greedy action: the action with the greatest current estimated value.
  2. Exploration chooses a nongreedy action to improve knowledge about its value.
  3. A single action selection cannot simultaneously exploit the current best-looking action and explore another action.
  4. Exploitation can maximize expected reward on the current step, while exploration may support greater total reward over many steps.
  5. Uncertainty matters because the current favorite may not truly be the best action.

Key Takeaways

  • A greedy action has the greatest current estimated value, so choosing it is exploitation.
  • A nongreedy action can be chosen for exploration when the goal is to improve knowledge about its value.
  • Immediate expected reward and long-run total reward can favor different actions.
  • Exploration can be worthwhile when current estimates are uncertain and learning may reveal a better action.
  • At one decision point, the agent must choose between exploiting current knowledge and exploring another action.