Concepts / Dyna-Q planning

Dyna-Q planning

Dyna-Q+ uses elapsed time since the last real trial as a signal for exploration.

  • Programming

Why Planning Needs Exploration

Dyna-Q planning uses a learned model to create simulated experiences. This can make an agent efficient because it can plan from what it already knows. However, that efficiency can create a blind spot: if planning repeatedly uses familiar transitions, the agent may spend too little effort checking actions that have been ignored for a long time. Dyna-Q+ addresses this problem by using elapsed time since the last real trial as an exploration signal.

Dyna-Q+ does not treat every modeled transition as equally attractive during planning. A state-action pair that has gone untested for longer receives a larger exploratory bonus.

Tracking Elapsed Time

For each state-action pair, Dyna-Q+ considers how much time has elapsed since that pair was last tried in real experience. Call this elapsed-time count τ. The count is not itself a reward. Instead, it supplies information about how neglected the pair has become. As τ becomes larger, the pair receives stronger encouragement during simulated planning.

more time passesmore time passesτ = 1bonus κτ = 4bonus 2κτ = 9bonus 3κ
How does the modeled reward change as more time passes since a state-action pair was tried?

Comparing two elapsed-time counts

Two state-action pairs are considered during planning. One has elapsed-time count τ = 1, and the other has elapsed-time count τ = 9. Which one receives the larger exploration bonus?

Identify the signal: Dyna-Q+ adds the bonus κ √τ to the modeled reward.

Compare the counts: The second pair has the larger elapsed-time count because 9 is greater than 1.

Compare the bonuses: The bonus for τ = 1 is κ, while the bonus for τ = 9 is 3κ. Therefore, the second pair receives the larger bonus.

The state-action pair that has gone untried longer receives greater encouragement during planning.

Adding the Exploration Bonus

The planning backup in Dyna-Q+ adds the bonus κ √τ to the modeled reward. Here, τ represents the elapsed time since the state-action pair was last tried, and κ controls the strength of this exploration encouragement. The important relationship is that the bonus grows with τ. Consequently, a transition that has remained untested for longer becomes more attractive in simulated planning.

add exploration bonusModeled rewardrAdjusted rewardr + κ √τ
How does Dyna-Q+ add extra encouragement to an action that has not been tried recently?

What do you think happens?

Two transitions have the same modeled reward, but one has a larger elapsed-time count. Which transition receives the larger effective reward during Dyna-Q+ planning?

  • The transition with the smaller elapsed-time count
  • Both receive the same effective reward
  • The transition with the larger elapsed-time count
Reveal answer

Answer: The transition with the larger elapsed-time count

The planning backup adds κ √τ. Since the square-root term is larger for the larger elapsed-time count, that transition receives the larger exploration bonus.

Including Never-Tried Actions

Dyna-Q+ extends exploration further than rewarding recently neglected actions. An action that has never been tried from a particular state is still allowed to enter the planning step. It is not excluded merely because no real transition has yet been observed for it.

For such a never-tried action, Dyna-Q+ begins with a deliberately simple model: the action is assumed to return the agent to the same state and produce reward zero. This initial same-state, zero-reward model gives the action a modeled transition that can participate in planning. The action can therefore receive an exploratory value rather than remaining completely absent from the planning process.

consider actioninitial modelenter planningState SUntried action ASame state Sreward 0Planning updateexploratory value
How can an action with no previous real experience enter the planning process and receive an exploratory value?

A missing action is not ignored

At state S, action A has never been tried. How can Dyna-Q+ include it in planning?

Create the initial model: The model represents action A as returning the agent to state S with reward zero.

Allow planning: Because the action now has an initial modeled transition, it can enter the planning step even without a previously observed real transition.

Apply exploration pressure: The action can receive an exploratory value through the Dyna-Q+ planning mechanism rather than being excluded as unknown.

A never-tried action participates in planning through the initial same-state, zero-reward model.

Planning Versus Exploration

Ordinary modeled planning uses the learned model to generate simulated experiences from transitions the agent has already represented. Dyna-Q+ keeps that model-based planning idea but adds an exploration signal based on elapsed time. The difference is not simply that Dyna-Q+ plans more; it changes which simulated experiences receive extra encouragement.

use modeltrack neglectadd bonususe modelModeled transitionlearned modelPlanning updatemodeled rewardModeled transitionlearned modelElapsed time τsince last trialPlanning updatereward plus κ √τ
What changes in the planning update when Dyna-Q+ uses time since the last trial as an exploration signal?
learn transitionmark last trialcompute signalprovide transitionadd encouragementReal trialstate-action pairLearned modelmodeled transitionExploration bonusκ √τPlanning backupadjusted modeled rewardElapsed time τsince last trial
How does information move from real experience and elapsed-time tracking into a modeled planning update?

Mistakes About Dyna-Q+

  • Treating the exploration bonus as an observed environmental reward

    The bonus is added during the planning backup to make a long-untried modeled transition more attractive.

    Fix: Treat κ √τ as an exploration signal used by planning.

  • Assuming only previously tried actions can be planned

    Dyna-Q+ allows a never-tried action to enter planning through an initial same-state, zero-reward model.

    Fix: Include the initial model for the never-tried action.

  • Giving the same bonus to every transition

    The bonus depends on τ, and a larger elapsed-time count produces a larger bonus.

    Fix: Compare elapsed-time counts when determining which transition receives more encouragement.

  • Confusing efficient planning with sufficient exploration

    Planning can become efficient while still creating a blind spot around actions that have been ignored for a long time.

    Fix: Use the elapsed-time signal to direct additional planning attention toward long-untried transitions.

Check Your Understanding

MEDIUM

Explain in your own words why Dyna-Q+ gives more encouragement to a state-action pair that has not been tried for a longer period. Then describe the initial model used when an action has never been tried from a state.

Hints
  • Start with the meaning of the elapsed-time count τ.
  • State how τ affects κ √τ.
  • Include both parts of the initial model for a never-tried action: the resulting state and the reward.
MEDIUM

Compare ordinary modeled planning with Dyna-Q+. Identify the extra information Dyna-Q+ uses and explain how that information changes the planning backup.

Hints
  • Both approaches use a learned model.
  • Focus on what elapsed time contributes.
  • Mention the exploration bonus rather than describing it as a real environmental reward.

Key Takeaways

  1. Dyna-Q+ uses elapsed time since the last real trial as an exploration signal.
  2. Its planning backup adds the bonus κ √τ to the modeled reward.
  3. The bonus becomes larger when a state-action pair has remained untested for longer.
  4. An action that has never been tried can still participate in planning through an initial same-state, zero-reward model.
  5. The distinctive exploration mechanism is the use of elapsed time to make neglected modeled experiences more attractive during planning.

Key Takeaways

  • Dyna-Q+ uses elapsed time since the last real trial to guide exploration during planning.
  • The exploration bonus κ √τ increases as the state-action pair remains untried for longer.
  • Never-tried actions are included through a same-state, zero-reward initial model.
  • Dyna-Q+ differs from ordinary modeled planning by adding an elapsed-time-based exploration signal to the planning backup.