Action Selection
Dyna-Q+ is Dyna-Q with an exploration bonus.
Why Action Selection Needs Exploration
An agent must choose actions using its current estimated values, but relying only on those estimates may not encourage enough exploration. Dyna-Q+ addresses this design issue by extending Dyna-Q with an exploration bonus. The bonus makes exploration more attractive, and the exercises in this topic examine where that bonus should enter the algorithm.
Dyna-Q+ is Dyna-Q with an exploration bonus. The important design question is whether the bonus should change estimated values during backups, affect action choice only, or be compared in both ways.
The Dyna-Q+ Information Path
To reason about Dyna-Q+, keep two quantities separate: the ordinary estimated value Q(S, a) and the additional exploration term. The source material identifies κ √ τ as the exploration bonus. When the bonus is used in backups, it changes estimated values. Those changed estimates can later influence action decisions. When the bonus is used only for selection, it changes the score used to choose an action without being described as a change to the ordinary estimated value.
Reading the Exploration Terms
κ √ τThe expression κ √ τ is the exploration bonus used in Dyna-Q+. It has two named parts in the source material: κ and τ. The expression is added to an estimated value to make exploration more attractive. The exercises use this expression to examine how an algorithm can encourage exploration rather than relying only on existing value estimates.
Q(S, a) + κ √ τSaExercise 8.4 places the bonus solely in action selection. Under that proposal, the selected action is the one that maximizes Q(S, a) + κ √ τSa. The first term is the ordinary estimated value for action a in state S. The second term is the exploration contribution associated with the state-action pair in the exercise. The selection rule therefore compares estimated value and exploration incentive together.
Backup Versus Selection
The location of the bonus matters because backups and action selection perform different jobs. A backup changes an estimated value used by the algorithm. If κ √ τ is incorporated into backups, the resulting changed estimates can influence later decisions. By contrast, the selection-only proposal leaves the ordinary estimated value Q(S, a) as the value term and chooses the action by maximizing Q(S, a) + κ √ τSa.
Comparing Two Candidate Actions
Suppose an exercise compares two actions in the same state using the selection-only rule. Action a1 has ordinary estimated value Q(S, a1), and action a2 has ordinary estimated value Q(S, a2). Each action also has its own exploration term in the expression Q(S, a) + κ √ τSa.
Keep the ordinary values separate: Start with Q(S, a1) and Q(S, a2). These are the ordinary estimated values and should not be confused with the exploration terms.
Add the corresponding exploration term: For each action, form its own score: Q(S, a1) + κ √ τSa1 and Q(S, a2) + κ √ τSa2.
Compare the scores: The selection-only rule chooses the action with the larger combined score. A lower ordinary value can be overcome by a larger exploration contribution if the combined score becomes larger.
Do not update the interpretation: This comparison describes action selection. It does not by itself say that the ordinary Q values were changed by a backup.
The selection-only approach combines the ordinary estimated value and the action-specific exploration contribution for the purpose of choosing an action, then selects the maximum.
Testing the Selection Rule
Design a gridworld experiment that compares Dyna-Q+ with a selection-only version. In your analysis, identify where κ √ τ is used in each version, record whether the bonus changes estimated values or only the action-selection score, and discuss strengths and weaknesses observed in the comparison.
Hints
- Trace the ordinary estimated value and the exploration term as separate quantities.
- For the selection-only version, apply the rule that chooses the action maximizing Q(S, a) + κ √ τSa.
- Compare the two approaches in the blocking and shortcut settings discussed by the source material.
What do you think happens?
If two actions have similar ordinary estimated values, what can make one action more attractive under the selection-only rule?
Reveal answer
Answer: A larger exploration contribution can make its combined score Q(S, a) + κ √ τSa larger.
The selection-only rule compares the ordinary estimated value together with the exploration term and chooses the action with the maximum combined score.
Common Reasoning Errors
Treating a bonus in backups and a bonus in action selection as the same operation.
Backups change estimated values, whereas selection-only use changes the score used to choose an action.
Fix:
State explicitly whether the bonus enters the value update or only the action-selection comparison.Ignoring the ordinary estimated value Q(S, a).
The selection-only rule maximizes the sum Q(S, a) + κ √ τSa, not the exploration term alone.
Fix:
Calculate or compare both parts of the combined score.Assuming that the selection-only rule automatically changes Q(S, a).
The exercise specifically changes the location of the bonus so that it is used solely during action selection.
Fix:
Keep the ordinary estimated value and the temporary selection score conceptually separate.Discussing performance without testing the alternate approach.
The exercise asks for an experiment that illustrates strengths and weaknesses.
Fix:
Compare the selection-only rule with the approach that uses the bonus in backups.
Key Takeaways
- Dyna-Q+ extends Dyna-Q by adding an exploration bonus.
- The bonus κ √ τ makes exploration more attractive by adding an exploration contribution to an estimated value or selection score.
- Using the bonus in backups can change estimated values and thereby influence later decisions.
- Using the bonus only during action selection chooses the action that maximizes Q(S, a) + κ √ τSa.
- A sound comparison must keep action-selection effects distinct from changes to estimated values and test both approaches experimentally.