Concepts / Action-Value Estimates

Action-Value Estimates

Greedy actions are those with the greatest current estimated value.

  • Machine Learning

The Choice at Each Step

When an agent must choose an action, it has two broad options. It can rely on what it currently believes about the available actions, or it can choose an action to learn more. Action-value estimates represent those current beliefs about how valuable the available actions appear to be.

The action with the greatest current estimated value is called the greedy action. Choosing it is exploitation: the agent uses its current knowledge to seek the best expected reward for the next step. Choosing a different action is exploration: the agent gathers information that may improve its estimates later.

Finding the Greedy Action

To identify a greedy action, compare the current estimated values of the available actions. The action or actions with the greatest estimate are greedy. If the agent selects one of those actions, it is exploiting its current knowledge. If it selects an action with a lower estimate, it is choosing a nongreedy action and exploring instead.

comparehighest estimatecompareAction Aestimated value 6Action Bgreatest current estimateAction Bestimated value 9Action Cestimated value 4
How do the estimated values for available actions determine which action is greedy?

Comparing three current estimates

Suppose an agent currently estimates the values of three available actions as follows: Action A is 6, Action B is 9, and Action C is 4. Which action is greedy, and what does selecting it represent?

Compare: Compare the three current estimates: 6, 9, and 4.

Identify the greatest estimate: Action B has the greatest current estimated value, 9.

Classify the selection: Action B is the greedy action. Selecting Action B is exploitation because the agent is using its current knowledge to seek the best expected reward for the next step.

Action B is greedy, and selecting it is exploitation.

Exploitation and Exploration

ChoiceAction selectedMain purposeTime horizon
ExploitationThe greedy actionUse current knowledge to seek the best expected one-step rewardImmediate
ExplorationA nongreedy actionGather information that may improve action-value estimatesLater selections

Exploitation can be the better choice when the immediate reward matters most. The agent selects the action that currently appears most valuable and does not spend this selection gathering information about another action.

Exploration can be the better choice when many future selections remain. A nongreedy action may reveal that the current estimates are incomplete or that another action is better than currently believed. That information can improve later choices and produce greater total reward over time.

aims atseeksExploitationgreedy actionImmediate rewarduse current knowledgeExplorationnongreedy actionNew informationimprove later estimates
What is the difference between selecting the action with the highest current estimate and trying another action to gather information?

Following the Decision Path

inspectidentify the greatest estimateseek immediate rewardgather informationAvailable actionscurrent estimatesCompare estimatesfind greatest valueSelection purposeuse knowledge or learnGreedy actionexploitationNongreedy actionexploration
Given current estimated values and an exploration rule, what decision path leads to a greedy or nongreedy action?

The reasoning has two stages. First, inspect the current estimates and identify which action is greedy. Second, decide whether this selection should use current knowledge or gather information. Choosing the greedy action serves exploitation. Choosing a nongreedy action serves exploration.

A single action selection cannot fully serve both purposes. Selecting the greedy action uses the current estimate to pursue immediate reward. Selecting a nongreedy action instead creates an opportunity to improve the estimate of that action. This is why exploitation and exploration are described as being in conflict at the moment of selection.

Choosing between immediate reward and information

An agent estimates Action A at 8 and Action B at 5. It must choose one action now, and it has many future selections remaining. What are the two possible reasons for choosing an action?

Identify the greedy action: Action A has the greater current estimate, so Action A is greedy.

Reason for choosing Action A: Choosing Action A is exploitation. The agent uses its current knowledge to seek the best expected one-step reward.

Reason for choosing Action B: Choosing Action B is exploration. The agent gives up the currently higher estimate in order to gather information that may improve later choices.

Compare the horizons: Action A is associated with the stronger immediate choice according to current knowledge. Action B may support greater total reward later if the information changes future estimates.

Action A is greedy and supports exploitation; Action B is nongreedy and supports exploration.

When Estimates Change

Exploration matters because an action's current estimate may not be the end of the story. Trying a nongreedy action gathers information. That information may improve the estimate of the tried action, which can affect which action has the greatest estimated value in a later selection.

A later change in the greedy action

Initially, an agent estimates Action A at 8 and Action B at 5. The agent explores Action B and gathers information that improves its estimate. How can the later decision differ from the initial decision?

Initial comparison: Action A has the greater current estimate, so Action A is initially greedy.

Exploration: The agent chooses Action B, a nongreedy action, to gather information rather than to pursue the best current estimate.

Estimate improvement: The information from trying Action B may improve the estimate associated with Action B.

Later comparison: When the agent compares the estimates again, Action B may have a higher estimate than before and may become the greedy action.

Exploration can change later estimates and can change which action is greedy.

Mistakes in Classifying Actions

  • Treating the greedy action as the action that is guaranteed to be best.

    Greedy refers to the greatest current estimate, not a guarantee about what is actually best.

    Fix: Describe the action as greedy according to the information currently available.

  • Calling every nongreedy choice a mistake.

    A nongreedy choice can be exploration intended to gather information for later decisions.

    Fix: Ask whether the action was selected to improve knowledge and potentially support greater future reward.

  • Assuming exploitation and exploration pursue the same horizon.

    Exploitation targets the best expected one-step reward, while exploration may improve estimates for later selections.

    Fix: State whether the reasoning emphasizes immediate reward or future information.

  • Saying that one selection simultaneously exploits and explores.

    A single selection chooses either the greedy action or a nongreedy action; those choices serve different purposes at that moment.

    Fix: Classify the selected action: greedy selection is exploitation, and nongreedy selection is exploration.

Practice the Reasoning

EASY

An agent currently estimates three actions as follows: Action A is 7, Action B is 3, and Action C is 5. First identify the greedy action. Then describe what the agent is doing if it selects Action A, and what it is doing if it selects Action C because it wants more information for future selections.

Hints
  • The greedy action has the greatest current estimated value.
  • Selecting the greedy action is exploitation.
  • Selecting a nongreedy action to gather information is exploration.

What do you think happens?

Using the practice estimates, which action is greedy, and how should selecting Action C be classified?

  • Action A is greedy; selecting Action C is exploration
  • Action B is greedy; selecting Action C is exploitation
  • Action C is greedy; selecting Action C is exploitation
Reveal answer

Answer: Action A is greedy; selecting Action C is exploration.

Action A has the greatest current estimated value, 7. Action C has a lower estimate, so selecting it is nongreedy exploration when the purpose is to gather information for future selections.

Key Takeaways

  1. A greedy action is an action with the greatest current estimated value.
  2. Selecting the greedy action is exploitation and aims at the best expected one-step reward.
  3. Selecting a nongreedy action is exploration and gathers information that may improve later estimates.
  4. Exploitation can favor immediate reward, while exploration can support greater total reward when many future selections remain.
  5. A single selection cannot simultaneously choose the greedy action for exploitation and a nongreedy action for exploration.

Key Takeaways

  • The greedy action is the action with the greatest current estimated value.
  • Exploitation uses current knowledge to seek the best expected immediate reward.
  • Exploration selects a nongreedy action to gather information for later decisions.
  • The better choice depends on whether the priority is immediate reward or improving future choices.
  • Exploration can change later estimates and therefore change which action is greedy.