Planning and Learning in Dyna
The Dyna Maze has 47 states, four directional actions, zero reward on ordinary transitions, and reward +1 when entering the goal.
Why Planning Matters
Dyna-Q addresses a practical question: after an agent has gathered experience, can it improve its policy more quickly by planning from that experience? The agent still interacts with the maze, but it also uses what it has already learned to generate additional learning opportunities. The central idea is that one real interaction can support both immediate learning and later simulated learning.
Planning is an additional way to improve a policy from existing experience; in this comparison, it does not replace interaction with the maze.
The Dyna Maze
The Dyna Maze contains 47 states. At each state, the agent has four directional actions available. Ordinary transitions provide reward 0. Entering the goal provides reward +1. The learning problem is therefore to discover actions that eventually lead to the goal and to improve the policy that reaches it.
The maze structure makes delayed information important. Most transitions have no immediate reward, while the useful reward appears when the agent enters the goal. A learning method must therefore use experience to improve actions that may occur several steps before the rewarding transition.
One Transition Two Uses
A real experience becomes reusable
Suppose the agent experiences one transition: it takes action a in state s, receives reward r, and arrives in state s'. How can this one event contribute to Dyna-Q twice?
Direct learning: The real transition is used immediately by the direct reinforcement-learning part of Dyna-Q, described in the source as one-step tabular Q-learning.
Model update: The experience also teaches the table-based model the recorded reward and next-state prediction associated with the experienced state-action pair.
Search control: During a later planning phase, search control selects a starting state and action from previously experienced state-action pairs.
Simulated experience: The model returns its recorded reward and next-state prediction for that selected pair, creating a simulated experience.
Planning update: The simulated experience is used by one-step tabular Q-planning, so the original real transition can influence learning again without requiring the agent to immediately repeat the environmental interaction.
A real transition supports an immediate direct-learning update and can later support one or more planning updates through the model.
The model does not replace the environment. It preserves information from real interaction so that the planning component can reuse that information later.
The Planning Budget n
In Dyna-Q, n specifies the number of planning steps performed per real step. It is useful to interpret n as a planning budget: n = 0 means no additional planning steps, while larger values add more simulated updates after each real interaction.
| Agent | Planning after each real step | Reported improvement |
|---|---|---|
| n = 0 | No planning; direct reinforcement learning only | About 25 episodes |
| n = 5 | Five planning steps | About 5 episodes |
| n = 50 | Fifty planning steps | About 3 episodes |
The comparison used the same kind of maze and varied how much planning occurred after each real step. The first episode was the same for every n value because the random-number-generator seed was held constant across the algorithms. That episode took about 1700 steps and was not shown in the reported curves.
How Planning Spreads Policy Information
Direct learning and planning differ in how quickly useful information can spread through the maze. With n = 0, the agent relies only on direct reinforcement learning. In the halfway-through-the-second-episode comparison, each episode adds only one more step to the learned policy, so only the final step has been learned at that point. With n = 50, the first episode still directly teaches only one step, but planning during the second episode develops a much larger policy while the agent is still wandering near the start. By the end of the second episode, the policy reaches almost back to the start state. By the end of the third episode, the source reports a complete optimal policy and perfect performance.
Planning spreads information because the model allows previously recorded transitions to be replayed as simulated experiences. Once useful information has been obtained near the goal, repeated planning updates can affect earlier parts of the learned policy while the agent is still exploring other parts of the maze.
Known Pairs and Search Control
Search control selects the starting states and actions for simulated experiences. Dyna-Q restricts these selections to state-action pairs that the agent has already experienced. That restriction matters because the table-based model has recorded reward and next-state information for experienced pairs. When search control selects one of those pairs, the model can return a simulated experience based on what was recorded.
Mistakes About Dyna-Q
Treating planning as a replacement for real interaction
The comparison describes planning as additional computation performed after real steps. The agent still interacts with the maze and builds its model from real experience.
Fix:
Think of n as an additional planning budget applied to what the agent has already experienced.Assuming n changes the maze
The experiment compares different amounts of planning on the same kind of task.
Fix:
n changes the number of planning steps performed after each real step, not the maze structure.Assuming the first episode proves that planning is faster
The first episode was the same for every n value because the random-number-generator seed was held constant across the algorithms.
Fix:
Compare what happens after the shared initial experience, when the agents use different planning budgets.Planning from arbitrary unexperienced pairs
The table-based model has no recorded reward and next-state prediction for that pair.
Fix:
Let search control select from previously experienced state-action pairs.Treating the reported n = 50 result as a universal rule
The reported learning speeds came from a particular maze experiment and particular parameter settings.
Fix:
Use the result to understand the effect of additional planning in this experiment, not as a universal guarantee.
Check Your Understanding
A Dyna-Q agent has just experienced a state-action-reward-next-state transition. Explain two different ways this transition can affect the agent: one immediate use and one later use. Then explain why a planning step should select a previously experienced state-action pair.
Hints
- Separate direct reinforcement learning from model-based planning.
- Identify what the model records from the real transition.
- Ask whether the model has a prediction for an unexperienced pair.
Compare n = 0 and n = 50 in the reported Dyna Maze experiment. Describe what happens after each real step and why the n = 50 agent can develop a larger policy while it is still wandering near the start.
Hints
- n = 0 performs no planning steps.
- n = 50 performs fifty planning steps after each real step.
- Planning reuses transitions recorded in the model.
Key Takeaways
- The Dyna Maze has 47 states, four directional actions, zero reward on ordinary transitions, and reward +1 when entering the goal.
- Dyna-Q combines one-step tabular Q-learning for direct reinforcement learning with random-sample one-step tabular Q-planning.
- The parameter n is the number of planning steps performed per real step; n = 0 is the nonplanning baseline.
- A real transition can update learning directly and also teach a model that later produces simulated experiences.
- Planning can spread useful policy information through the maze faster than direct learning alone, as shown by the reported improvement for larger n values.
Key Takeaways
- Dyna-Q learns from real interaction and also plans from a model built from that interaction.
- The planning budget n controls how many simulated planning updates occur after each real step.
- In the reported Dyna Maze experiment, larger planning values led to faster improvement: about 25 episodes for n = 0, about 5 for n = 5, and about 3 for n = 50.
- Planning reuses experienced state-action pairs so that useful information can spread through the learned policy while the agent continues interacting with the maze.