Concepts / Exploration-Exploitation Trade-off

Exploration-Exploitation Trade-off

Dyna-Q+ is Dyna-Q with an exploration bonus.

  • Programming

One Choice, Two Goals

At every time step, an action-selection method faces a conflict. It can choose the action that currently appears best and seek the greatest expected reward on this step, or it can choose another action to improve its estimate of that action's value. The first choice is exploitation. The second is exploration. Since one action selection cannot do both at once, deciding between them creates the exploration-exploitation trade-off.

usesusesGreedy actiongreatest estimated valueExploitationimmediate expected rewardNongreedy actionimproves knowledgeExplorationinformation about value
How does the selected action differ when the agent follows its current estimate versus when it investigates another action?

Exploitation uses the action that currently has the greatest estimated value. Exploration chooses a nongreedy action to improve knowledge about its value.

Dyna-Q+ Adds Information-Seeking Value

Dyna-Q+ is Dyna-Q with an exploration bonus. The bonus is added to an estimated value to make exploration more attractive. In Dyna-Q+, the exploration bonus is written as κ √ τ. This gives the algorithm a mechanism for encouraging exploration instead of relying only on its existing value estimates.

tracksenterschangesinfluencesLearned modelsimulated experienceκ √ τexploration bonusSimulated backupbonus incorporatedEstimated valuesupdated estimatesLater decisionsinfluenced by estimates
How does Dyna-Q+ add an exploration bonus as information flows from the learned model into simulated backups and updated action values?

To reason about Dyna-Q+, keep two quantities separate: the ordinary estimated value and the additional exploration term. The expression κ √ τ is the additional term. When the bonus is incorporated into backups, it changes estimated values of states and actions. Those changed estimates can influence later decisions.

Reading the Exploration Expressions

κ √ τ

Q(S, a) + κ √ τSa

first termaddsmultipliesapplies toQ(S, a)ordinary estimated value+combineκexploration coefficient√square rootτSatime since last trial
How do the estimated value, exploration coefficient, and time-since-last-trial quantity combine during selection?

Comparing Two Selection Scores

Suppose two actions have different ordinary estimated values and different exploration terms. Use the selection-only expression to compare them.

Action A: Generated illustration: its ordinary estimated value is 5 and its exploration term is 1, giving a combined selection score of 6.

Action B: Generated illustration: its ordinary estimated value is 4 and its exploration term is 3, giving a combined selection score of 7.

Compare: The selection-only rule compares the combined scores, not the ordinary estimated values alone.

Generated illustration: Action B would be selected because 7 is greater than 6, even though Action A has the greater ordinary estimated value.

Two Locations for the Bonus

The location of the exploration bonus matters. If the bonus is used in backups, it changes the estimated values. Those changed estimates remain available to influence later decisions. If the bonus is used only during action selection, the ordinary estimated values are not the quantity being changed by that bonus; instead, the agent temporarily combines Q(S, a) with κ √ τSa while deciding which action to choose.

enters backupmaximized over actionsκ √ τbonusEstimated valueschanged by backupQ(S, a) + κ √ τSaselection scoreSelected actionmaximal combined score
What changes when the exploration bonus is added inside the backup calculation compared with adding it only while choosing an action?
QuestionBonus in backupsBonus only during selection
Where is κ √ τ used?Inside the backup calculationWhile choosing an action
What is directly affected?Estimated valuesThe action-selection comparison
Rule highlighted in the exercisesBonus changes estimates that can influence later decisionsChoose the action maximizing Q(S, a) + κ √ τSa

Short-Term Cost, Long-Term Gain

Exploitation is appropriate when the goal is to maximize expected reward on the one step. However, the action that currently looks best may not truly be the best action. Exploration accepts a lower reward in the short run when it can improve knowledge about another action. If that investigation reveals a better action, the better action can then be exploited repeatedly, producing greater total reward over the long run.

investigatesmay enableExplorelower short-run rewardImprove estimatelearn another valueExploit better actiongreater later reward
How can choosing a lower immediate reward lead to greater total reward over multiple time steps?

A Deliberate Short-Term Sacrifice

An action currently appears best, but another action has an uncertain estimate. Explain why trying the uncertain action can be reasonable.

Immediate view: Generated illustration: choosing the currently best-looking action favors exploitation because it appears to provide the greatest expected reward on this step.

Information view: The uncertain action may be worth trying because the trial can improve the estimate of its value.

Long-run view: If the trial reveals that the uncertain action is better, the agent can exploit that better action repeatedly on later steps.

Generated illustration: exploration can be worthwhile when its possible information gain may lead to greater total reward over many time steps.

Uncertainty Changes the Calculation

An action value is an estimate, so the action that currently has the greatest estimate may not truly be the best action. Exploration is especially relevant when another action's value is not well known. The exploration bonus provides a way to make that investigation more attractive instead of allowing the current estimates alone to determine every choice.

compareadd exploration incentiveAction Ahigher current estimateAction Aordinary estimateAction Buncertain estimateAction Bbonus increases appeal
How can uncertainty in action-value estimates make an apparently inferior action worth trying?

The exploration bonus does not claim that a nongreedy action is already better according to the ordinary estimate. It makes exploration more attractive because learning more about that action may improve later decisions.

Common Reasoning Errors

  • Treating exploitation as any action with a positive reward.

    Exploitation is defined by choosing the action with the greatest current estimated value, not simply by receiving a positive reward.

    Fix: Ask which action currently has the greatest estimated value.

  • Assuming that exploration and exploitation happen simultaneously in one action selection.

    A single action selection cannot choose both alternatives at once.

    Fix: Classify the selected action: greedy selection is exploitation, while nongreedy selection for improving knowledge is exploration.

  • Ignoring where the bonus is applied.

    A bonus in backups changes estimated values, whereas a selection-only bonus changes the comparison used to choose an action.

    Fix: First identify whether κ √ τ appears in a backup or only in the action-selection expression.

  • Comparing only Q(S, a) when the rule is selection-only.

    The selection-only rule maximizes the combined expression over actions.

    Fix: Compute or compare the ordinary estimate plus its exploration term.

  • Assuming the immediate best-looking action must maximize total reward.

    Exploration can improve knowledge and may reveal an action that produces greater reward over many later steps.

    Fix: Separate the one-step objective from the long-run effect of improved estimates.

Gridworld Investigation

Exercise 8.4 changes the location of the bonus: instead of using κ √ τ in backups, it proposes using the bonus solely during action selection. The purpose is not merely to write the rule. The exercise asks for a gridworld experiment that tests and illustrates the strengths and weaknesses of the selection-only approach.

  1. State which approach is being tested: bonus in backups or bonus only during selection.
  2. For the selection-only approach, use Q(S, a) + κ √ τSa as the action-selection score.
  3. Keep ordinary estimated values separate from the additional exploration term.
  4. Compare the selection-only rule with the approach in which the bonus is used in backups.
  5. Interpret the results in terms of action choice, changed estimated values, and performance over the task.
MEDIUM

Explain why an action with a lower ordinary Q(S, a) can still be selected by the selection-only rule. Then explain how the answer would differ if the exploration bonus were incorporated into backups.

Hints
  • Compare the ordinary estimated value with the combined selection score.
  • Remember that backup use changes estimated values, while selection-only use changes the action comparison.
  • Discuss both immediate action choice and later decisions.

What do you think happens?

If two actions have ordinary estimated values of 5 and 4, but their exploration terms are 1 and 3 respectively, which action has the larger selection score?

  • The action with ordinary estimate 5
  • The action with ordinary estimate 4
  • They have equal selection scores
Reveal answer

Answer: The action with ordinary estimate 4 has the larger selection score.

The combined scores are 5 + 1 = 6 and 4 + 3 = 7. The selection-only rule maximizes the combined expression, so the second action is selected in this generated illustration.

Performance Evidence

The blocking and shortcut comparison reports better performance for Dyna-Q+ in both phases. This result supports studying the exploration bonus as a meaningful extension of Dyna-Q, while still leaving an important design question: should exploration be encouraged by changing estimated values through backups, or by leaving those estimates alone and adding a bonus only during action selection?

Key Takeaways

  1. Dyna-Q+ extends Dyna-Q with an exploration bonus written as κ √ τ. Exploitation selects the action with the greatest current estimated value, while exploration selects a nongreedy action to improve knowledge. A single action selection cannot perform both roles simultaneously. The location of the bonus matters: using it in backups changes estimated values, while using it only during selection changes the score used to choose an action. In the selection-only rule, the chosen action maximizes Q(S, a) + κ √ τSa. Exploration may accept lower immediate reward when improved knowledge can lead to greater total reward over many time steps.

Key Takeaways

  • Dyna-Q+ is Dyna-Q with an exploration bonus.
  • The expression κ √ τ represents the additional exploration term.
  • Adding the bonus in backups changes estimated values; adding it only during selection changes the action-selection comparison.
  • The selection-only rule chooses the action maximizing Q(S, a) + κ √ τSa.
  • Exploration can sacrifice immediate reward to improve knowledge and potentially increase total reward over many time steps.