Exploration-Exploitation Trade-off
Dyna-Q+ is Dyna-Q with an exploration bonus.
One Choice, Two Goals
At every time step, an action-selection method faces a conflict. It can choose the action that currently appears best and seek the greatest expected reward on this step, or it can choose another action to improve its estimate of that action's value. The first choice is exploitation. The second is exploration. Since one action selection cannot do both at once, deciding between them creates the exploration-exploitation trade-off.
Exploitation uses the action that currently has the greatest estimated value. Exploration chooses a nongreedy action to improve knowledge about its value.
Dyna-Q+ Adds Information-Seeking Value
Dyna-Q+ is Dyna-Q with an exploration bonus. The bonus is added to an estimated value to make exploration more attractive. In Dyna-Q+, the exploration bonus is written as κ √ τ. This gives the algorithm a mechanism for encouraging exploration instead of relying only on its existing value estimates.
To reason about Dyna-Q+, keep two quantities separate: the ordinary estimated value and the additional exploration term. The expression κ √ τ is the additional term. When the bonus is incorporated into backups, it changes estimated values of states and actions. Those changed estimates can influence later decisions.
Reading the Exploration Expressions
κ √ τ
Q(S, a) + κ √ τSa
Comparing Two Selection Scores
Suppose two actions have different ordinary estimated values and different exploration terms. Use the selection-only expression to compare them.
Action A: Generated illustration: its ordinary estimated value is 5 and its exploration term is 1, giving a combined selection score of 6.
Action B: Generated illustration: its ordinary estimated value is 4 and its exploration term is 3, giving a combined selection score of 7.
Compare: The selection-only rule compares the combined scores, not the ordinary estimated values alone.
Generated illustration: Action B would be selected because 7 is greater than 6, even though Action A has the greater ordinary estimated value.
Two Locations for the Bonus
The location of the exploration bonus matters. If the bonus is used in backups, it changes the estimated values. Those changed estimates remain available to influence later decisions. If the bonus is used only during action selection, the ordinary estimated values are not the quantity being changed by that bonus; instead, the agent temporarily combines Q(S, a) with κ √ τSa while deciding which action to choose.
| Question | Bonus in backups | Bonus only during selection |
|---|---|---|
| Where is κ √ τ used? | Inside the backup calculation | While choosing an action |
| What is directly affected? | Estimated values | The action-selection comparison |
| Rule highlighted in the exercises | Bonus changes estimates that can influence later decisions | Choose the action maximizing Q(S, a) + κ √ τSa |
Short-Term Cost, Long-Term Gain
Exploitation is appropriate when the goal is to maximize expected reward on the one step. However, the action that currently looks best may not truly be the best action. Exploration accepts a lower reward in the short run when it can improve knowledge about another action. If that investigation reveals a better action, the better action can then be exploited repeatedly, producing greater total reward over the long run.
A Deliberate Short-Term Sacrifice
An action currently appears best, but another action has an uncertain estimate. Explain why trying the uncertain action can be reasonable.
Immediate view: Generated illustration: choosing the currently best-looking action favors exploitation because it appears to provide the greatest expected reward on this step.
Information view: The uncertain action may be worth trying because the trial can improve the estimate of its value.
Long-run view: If the trial reveals that the uncertain action is better, the agent can exploit that better action repeatedly on later steps.
Generated illustration: exploration can be worthwhile when its possible information gain may lead to greater total reward over many time steps.
Uncertainty Changes the Calculation
An action value is an estimate, so the action that currently has the greatest estimate may not truly be the best action. Exploration is especially relevant when another action's value is not well known. The exploration bonus provides a way to make that investigation more attractive instead of allowing the current estimates alone to determine every choice.
The exploration bonus does not claim that a nongreedy action is already better according to the ordinary estimate. It makes exploration more attractive because learning more about that action may improve later decisions.
Common Reasoning Errors
Treating exploitation as any action with a positive reward.
Exploitation is defined by choosing the action with the greatest current estimated value, not simply by receiving a positive reward.
Fix:
Ask which action currently has the greatest estimated value.Assuming that exploration and exploitation happen simultaneously in one action selection.
A single action selection cannot choose both alternatives at once.
Fix:
Classify the selected action: greedy selection is exploitation, while nongreedy selection for improving knowledge is exploration.Ignoring where the bonus is applied.
A bonus in backups changes estimated values, whereas a selection-only bonus changes the comparison used to choose an action.
Fix:
First identify whether κ √ τ appears in a backup or only in the action-selection expression.Comparing only Q(S, a) when the rule is selection-only.
The selection-only rule maximizes the combined expression over actions.
Fix:
Compute or compare the ordinary estimate plus its exploration term.Assuming the immediate best-looking action must maximize total reward.
Exploration can improve knowledge and may reveal an action that produces greater reward over many later steps.
Fix:
Separate the one-step objective from the long-run effect of improved estimates.
Gridworld Investigation
Exercise 8.4 changes the location of the bonus: instead of using κ √ τ in backups, it proposes using the bonus solely during action selection. The purpose is not merely to write the rule. The exercise asks for a gridworld experiment that tests and illustrates the strengths and weaknesses of the selection-only approach.
- State which approach is being tested: bonus in backups or bonus only during selection.
- For the selection-only approach, use Q(S, a) + κ √ τSa as the action-selection score.
- Keep ordinary estimated values separate from the additional exploration term.
- Compare the selection-only rule with the approach in which the bonus is used in backups.
- Interpret the results in terms of action choice, changed estimated values, and performance over the task.
Explain why an action with a lower ordinary Q(S, a) can still be selected by the selection-only rule. Then explain how the answer would differ if the exploration bonus were incorporated into backups.
Hints
- Compare the ordinary estimated value with the combined selection score.
- Remember that backup use changes estimated values, while selection-only use changes the action comparison.
- Discuss both immediate action choice and later decisions.
What do you think happens?
If two actions have ordinary estimated values of 5 and 4, but their exploration terms are 1 and 3 respectively, which action has the larger selection score?
Reveal answer
Answer: The action with ordinary estimate 4 has the larger selection score.
The combined scores are 5 + 1 = 6 and 4 + 3 = 7. The selection-only rule maximizes the combined expression, so the second action is selected in this generated illustration.
Performance Evidence
The blocking and shortcut comparison reports better performance for Dyna-Q+ in both phases. This result supports studying the exploration bonus as a meaningful extension of Dyna-Q, while still leaving an important design question: should exploration be encouraged by changing estimated values through backups, or by leaving those estimates alone and adding a bonus only during action selection?
Key Takeaways
- Dyna-Q+ extends Dyna-Q with an exploration bonus written as κ √ τ. Exploitation selects the action with the greatest current estimated value, while exploration selects a nongreedy action to improve knowledge. A single action selection cannot perform both roles simultaneously. The location of the bonus matters: using it in backups changes estimated values, while using it only during selection changes the score used to choose an action. In the selection-only rule, the chosen action maximizes Q(S, a) + κ √ τSa. Exploration may accept lower immediate reward when improved knowledge can lead to greater total reward over many time steps.
Key Takeaways
- Dyna-Q+ is Dyna-Q with an exploration bonus.
- The expression κ √ τ represents the additional exploration term.
- Adding the bonus in backups changes estimated values; adding it only during selection changes the action-selection comparison.
- The selection-only rule chooses the action maximizing Q(S, a) + κ √ τSa.
- Exploration can sacrifice immediate reward to improve knowledge and potentially increase total reward over many time steps.