Concepts / Action Selection

Action Selection

Dyna-Q+ is Dyna-Q with an exploration bonus.

  • Programming

Why Action Selection Needs Exploration

An agent must choose actions using its current estimated values, but relying only on those estimates may not encourage enough exploration. Dyna-Q+ addresses this design issue by extending Dyna-Q with an exploration bonus. The bonus makes exploration more attractive, and the exercises in this topic examine where that bonus should enter the algorithm.

Dyna-Q+ is Dyna-Q with an exploration bonus. The important design question is whether the bonus should change estimated values during backups, affect action choice only, or be compared in both ways.

The Dyna-Q+ Information Path

To reason about Dyna-Q+, keep two quantities separate: the ordinary estimated value Q(S, a) and the additional exploration term. The source material identifies κ √ τ as the exploration bonus. When the bonus is used in backups, it changes estimated values. Those changed estimates can later influence action decisions. When the bonus is used only for selection, it changes the score used to choose an action without being described as a change to the ordinary estimated value.

updatessupplies informationcan change estimatessupplies Q(S, a)can adjust selection scoreExperienceobserved interactionModelstored informationPlanningbackupsEstimated valuesQ(S, a)Action selectionchoose an actionκ √ τexploration bonus
Where does the exploration bonus enter the Dyna-Q+ cycle, and how can it influence later action selection?

Reading the Exploration Terms

κ √ τ

The expression κ √ τ is the exploration bonus used in Dyna-Q+. It has two named parts in the source material: κ and τ. The expression is added to an estimated value to make exploration more attractive. The exercises use this expression to examine how an algorithm can encourage exploration rather than relying only on existing value estimates.

Q(S, a) + κ √ τSa

Exercise 8.4 places the bonus solely in action selection. Under that proposal, the selected action is the one that maximizes Q(S, a) + κ √ τSa. The first term is the ordinary estimated value for action a in state S. The second term is the exploration contribution associated with the state-action pair in the exercise. The selection rule therefore compares estimated value and exploration incentive together.

Q(S, a)ordinary estimated value+combine termsκbonus scale√ τSaexploration contribution
What does each part of Q(S, a) + κ √ τSa contribute to the score used for selecting an action?

Backup Versus Selection

The location of the bonus matters because backups and action selection perform different jobs. A backup changes an estimated value used by the algorithm. If κ √ τ is incorporated into backups, the resulting changed estimates can influence later decisions. By contrast, the selection-only proposal leaves the ordinary estimated value Q(S, a) as the value term and chooses the action by maximizing Q(S, a) + κ √ τSa.

changesformsBonus in backupsestimated values can changeChanged estimatesinfluence later decisionsBonus in selectionselection score changesQ(S, a) + κ √ τSachoose the maximum
What changes when the exploration bonus is added during backups compared with adding it only to the action-selection score?

Comparing Two Candidate Actions

Suppose an exercise compares two actions in the same state using the selection-only rule. Action a1 has ordinary estimated value Q(S, a1), and action a2 has ordinary estimated value Q(S, a2). Each action also has its own exploration term in the expression Q(S, a) + κ √ τSa.

Keep the ordinary values separate: Start with Q(S, a1) and Q(S, a2). These are the ordinary estimated values and should not be confused with the exploration terms.

Add the corresponding exploration term: For each action, form its own score: Q(S, a1) + κ √ τSa1 and Q(S, a2) + κ √ τSa2.

Compare the scores: The selection-only rule chooses the action with the larger combined score. A lower ordinary value can be overcome by a larger exploration contribution if the combined score becomes larger.

Do not update the interpretation: This comparison describes action selection. It does not by itself say that the ordinary Q values were changed by a backup.

The selection-only approach combines the ordinary estimated value and the action-specific exploration contribution for the purpose of choosing an action, then selects the maximum.

Testing the Selection Rule

MEDIUM

Design a gridworld experiment that compares Dyna-Q+ with a selection-only version. In your analysis, identify where κ √ τ is used in each version, record whether the bonus changes estimated values or only the action-selection score, and discuss strengths and weaknesses observed in the comparison.

Hints
  • Trace the ordinary estimated value and the exploration term as separate quantities.
  • For the selection-only version, apply the rule that chooses the action maximizing Q(S, a) + κ √ τSa.
  • Compare the two approaches in the blocking and shortcut settings discussed by the source material.

What do you think happens?

If two actions have similar ordinary estimated values, what can make one action more attractive under the selection-only rule?

Reveal answer

Answer: A larger exploration contribution can make its combined score Q(S, a) + κ √ τSa larger.

The selection-only rule compares the ordinary estimated value together with the exploration term and chooses the action with the maximum combined score.

Common Reasoning Errors

  • Treating a bonus in backups and a bonus in action selection as the same operation.

    Backups change estimated values, whereas selection-only use changes the score used to choose an action.

    Fix: State explicitly whether the bonus enters the value update or only the action-selection comparison.

  • Ignoring the ordinary estimated value Q(S, a).

    The selection-only rule maximizes the sum Q(S, a) + κ √ τSa, not the exploration term alone.

    Fix: Calculate or compare both parts of the combined score.

  • Assuming that the selection-only rule automatically changes Q(S, a).

    The exercise specifically changes the location of the bonus so that it is used solely during action selection.

    Fix: Keep the ordinary estimated value and the temporary selection score conceptually separate.

  • Discussing performance without testing the alternate approach.

    The exercise asks for an experiment that illustrates strengths and weaknesses.

    Fix: Compare the selection-only rule with the approach that uses the bonus in backups.

Key Takeaways

  • Dyna-Q+ extends Dyna-Q by adding an exploration bonus.
  • The bonus κ √ τ makes exploration more attractive by adding an exploration contribution to an estimated value or selection score.
  • Using the bonus in backups can change estimated values and thereby influence later decisions.
  • Using the bonus only during action selection chooses the action that maximizes Q(S, a) + κ √ τSa.
  • A sound comparison must keep action-selection effects distinct from changes to estimated values and test both approaches experimentally.