Dyna-Q planning
Dyna-Q+ uses elapsed time since the last real trial as a signal for exploration.
Why Planning Needs Exploration
Dyna-Q planning uses a learned model to create simulated experiences. This can make an agent efficient because it can plan from what it already knows. However, that efficiency can create a blind spot: if planning repeatedly uses familiar transitions, the agent may spend too little effort checking actions that have been ignored for a long time. Dyna-Q+ addresses this problem by using elapsed time since the last real trial as an exploration signal.
Dyna-Q+ does not treat every modeled transition as equally attractive during planning. A state-action pair that has gone untested for longer receives a larger exploratory bonus.
Tracking Elapsed Time
For each state-action pair, Dyna-Q+ considers how much time has elapsed since that pair was last tried in real experience. Call this elapsed-time count τ. The count is not itself a reward. Instead, it supplies information about how neglected the pair has become. As τ becomes larger, the pair receives stronger encouragement during simulated planning.
Comparing two elapsed-time counts
Two state-action pairs are considered during planning. One has elapsed-time count τ = 1, and the other has elapsed-time count τ = 9. Which one receives the larger exploration bonus?
Identify the signal: Dyna-Q+ adds the bonus κ √τ to the modeled reward.
Compare the counts: The second pair has the larger elapsed-time count because 9 is greater than 1.
Compare the bonuses: The bonus for τ = 1 is κ, while the bonus for τ = 9 is 3κ. Therefore, the second pair receives the larger bonus.
The state-action pair that has gone untried longer receives greater encouragement during planning.
Adding the Exploration Bonus
The planning backup in Dyna-Q+ adds the bonus κ √τ to the modeled reward. Here, τ represents the elapsed time since the state-action pair was last tried, and κ controls the strength of this exploration encouragement. The important relationship is that the bonus grows with τ. Consequently, a transition that has remained untested for longer becomes more attractive in simulated planning.
What do you think happens?
Two transitions have the same modeled reward, but one has a larger elapsed-time count. Which transition receives the larger effective reward during Dyna-Q+ planning?
Reveal answer
Answer: The transition with the larger elapsed-time count
The planning backup adds κ √τ. Since the square-root term is larger for the larger elapsed-time count, that transition receives the larger exploration bonus.
Including Never-Tried Actions
Dyna-Q+ extends exploration further than rewarding recently neglected actions. An action that has never been tried from a particular state is still allowed to enter the planning step. It is not excluded merely because no real transition has yet been observed for it.
For such a never-tried action, Dyna-Q+ begins with a deliberately simple model: the action is assumed to return the agent to the same state and produce reward zero. This initial same-state, zero-reward model gives the action a modeled transition that can participate in planning. The action can therefore receive an exploratory value rather than remaining completely absent from the planning process.
A missing action is not ignored
At state S, action A has never been tried. How can Dyna-Q+ include it in planning?
Create the initial model: The model represents action A as returning the agent to state S with reward zero.
Allow planning: Because the action now has an initial modeled transition, it can enter the planning step even without a previously observed real transition.
Apply exploration pressure: The action can receive an exploratory value through the Dyna-Q+ planning mechanism rather than being excluded as unknown.
A never-tried action participates in planning through the initial same-state, zero-reward model.
Planning Versus Exploration
Ordinary modeled planning uses the learned model to generate simulated experiences from transitions the agent has already represented. Dyna-Q+ keeps that model-based planning idea but adds an exploration signal based on elapsed time. The difference is not simply that Dyna-Q+ plans more; it changes which simulated experiences receive extra encouragement.
Mistakes About Dyna-Q+
Treating the exploration bonus as an observed environmental reward
The bonus is added during the planning backup to make a long-untried modeled transition more attractive.
Fix:
Treat κ √τ as an exploration signal used by planning.Assuming only previously tried actions can be planned
Dyna-Q+ allows a never-tried action to enter planning through an initial same-state, zero-reward model.
Fix:
Include the initial model for the never-tried action.Giving the same bonus to every transition
The bonus depends on τ, and a larger elapsed-time count produces a larger bonus.
Fix:
Compare elapsed-time counts when determining which transition receives more encouragement.Confusing efficient planning with sufficient exploration
Planning can become efficient while still creating a blind spot around actions that have been ignored for a long time.
Fix:
Use the elapsed-time signal to direct additional planning attention toward long-untried transitions.
Check Your Understanding
Explain in your own words why Dyna-Q+ gives more encouragement to a state-action pair that has not been tried for a longer period. Then describe the initial model used when an action has never been tried from a state.
Hints
- Start with the meaning of the elapsed-time count τ.
- State how τ affects κ √τ.
- Include both parts of the initial model for a never-tried action: the resulting state and the reward.
Compare ordinary modeled planning with Dyna-Q+. Identify the extra information Dyna-Q+ uses and explain how that information changes the planning backup.
Hints
- Both approaches use a learned model.
- Focus on what elapsed time contributes.
- Mention the exploration bonus rather than describing it as a real environmental reward.
Key Takeaways
- Dyna-Q+ uses elapsed time since the last real trial as an exploration signal.
- Its planning backup adds the bonus κ √τ to the modeled reward.
- The bonus becomes larger when a state-action pair has remained untested for longer.
- An action that has never been tried can still participate in planning through an initial same-state, zero-reward model.
- The distinctive exploration mechanism is the use of elapsed time to make neglected modeled experiences more attractive during planning.
Key Takeaways
- Dyna-Q+ uses elapsed time since the last real trial to guide exploration during planning.
- The exploration bonus κ √τ increases as the state-action pair remains untried for longer.
- Never-tried actions are included through a same-state, zero-reward initial model.
- Dyna-Q+ differs from ordinary modeled planning by adding an elapsed-time-based exploration signal to the planning backup.