Concepts / One-Step-Ahead Search

One-Step-Ahead Search

q* turns optimal action selection into comparing values for the current state's available actions.

  • Programming

Choosing Without Repeating the Search

Whenever an agent is in a state, it must choose an action. One possible method is to examine an action, consider what may happen next, and use those possible outcomes to judge the action. This is a one-step-ahead search. The optimal action-value function q* changes when that work happens: instead of repeating the look-ahead process at decision time, the agent can compare the q* values already available for the current state's actions.

The central operation is simple: in the current state, select any available action that maximizes q*(s, a).

Reading a State-Action Value

The expression q*(s, a) associates a value with taking action a from state s. The function stores the results of one-step-ahead searches as immediately available state-action values. Therefore, when the agent reaches s, it can inspect the values associated with the available actions instead of carrying out the look-ahead process again.

available actionavailable actionavailable actionState scurrent stateAction a1q*(s, a1)Action a2q*(s, a2)Action a3q*(s, a3)
What value does q*(s, a) associate with each action available from the current state?

The diagram represents q* as a collection of values attached to the actions available from one state. The agent's selection task is to compare those values. The value does not require the agent to discover the successor state during this selection step, because the result needed for comparison is already stored in q*.

From Look-Ahead to Lookup

startlook aheadevaluate outcomeslook upcomparechoose maximumCurrent statebeforeCurrent stateafterConsider actionbeforeq* valuesafterSuccessor statesbeforeCompare actionsafterJudge actionbeforeOptimal actionafter
What changes when action selection uses stored q* values instead of predicting successor states and evaluating them?

In the search approach, the agent considers an action, examines what may happen next, and uses the possible outcomes to judge that action. With q*, the one-step-ahead results are already available as state-action values. Selection becomes a lookup-and-compare process: identify the current state, inspect q* for its available actions, and choose an action with the largest value.

Selecting from Equal Maximum Values

Two Actions Share the Best Value

In a state, action a1 has q* value 6 and action a2 also has q* value 6. Which action can the agent select as an optimal action?

Compare the values: The available actions have equal q* values: both a1 and a2 have value 6.

Find the maximum: The maximum value among the available actions is 6.

Apply the maximizing choice: Any action that maximizes q*(s, a) can be selected as an optimal action.

The agent can select either a1 or a2. No additional one-step-ahead search is needed because q* already supplies the values used for comparison.

maximummaximuma1q* = 6Optimal actiona1 or a2a2q* = 6
Given a state and several available actions, which action is selected when the agent chooses the largest q* value?

When several actions share the maximum q* value, each of those actions is an optimal choice according to the maximizing rule.

Common Selection Mistakes

  • Choosing an action without comparing the available q* values.

    The optimal action is obtained by choosing an action that maximizes q*(s, a).

    Fix: Compare the q* values for all available actions before selecting one.

  • Assuming that equal maximum values make the choice invalid.

    Both actions maximize the available q* values.

    Fix: Select either action that has the maximum value.

  • Repeating a one-step-ahead search even though q* already provides the comparison values.

    q* stores the results of one-step-ahead searches as immediately available state-action values.

    Fix: Use q* to compare the current state's available actions directly.

Apply the Maximizing Rule

EASY

An agent is in state s. Three actions are available: a1 has q*(s, a1) equal to 3, a2 has q*(s, a2) equal to 8, and a3 has q*(s, a3) equal to 5. Which action should the agent select as an optimal action?

Hints
  • Compare the three q* values.
  • Select an action with the largest value.

For the practice values, a2 is the optimal selection because its q* value is the largest of the three. The example illustrates the same maximizing operation described by the source concept.

The Lookup-and-Compare Pattern

  1. q*(s, a) provides a value for taking action a from state s.
  2. q* stores the results of one-step-ahead searches as immediately available state-action values.
  3. At a state, the agent can select any available action that maximizes q*(s, a).
  4. Because q* already supplies the values for comparison, the agent does not need successor-state information or environment dynamics during this selection.
  5. If multiple actions share the maximum value, any one of them can be selected as an optimal action.

Key Takeaways

  • q*(s, a) associates a stored value with an action taken from a particular state.
  • One-step-ahead search evaluates what may happen next; q* makes the resulting values immediately available for selection.
  • Optimal action selection using q* is a lookup-and-compare process.
  • The agent chooses any available action with the largest q* value and does not need successor-state information for that selection.
  • Ties are allowed: if multiple actions share the maximum value, any of them is optimal.