Concepts / Planning and Learning in Dyna

Planning and Learning in Dyna

The Dyna Maze has 47 states, four directional actions, zero reward on ordinary transitions, and reward +1 when entering the goal.

  • Programming

Why Planning Matters

Dyna-Q addresses a practical question: after an agent has gathered experience, can it improve its policy more quickly by planning from that experience? The agent still interacts with the maze, but it also uses what it has already learned to generate additional learning opportunities. The central idea is that one real interaction can support both immediate learning and later simulated learning.

Planning is an additional way to improve a policy from existing experience; in this comparison, it does not replace interaction with the maze.

The Dyna Maze

The Dyna Maze contains 47 states. At each state, the agent has four directional actions available. Ordinary transitions provide reward 0. Entering the goal provides reward +1. The learning problem is therefore to discover actions that eventually lead to the goal and to improve the policy that reaches it.

ordinary actionentering goalMaze state47-state taskOrdinary next statereward 0Goalreward +1 on entry
How do ordinary movements differ from the transition that enters the goal?

The maze structure makes delayed information important. Most transitions have no immediate reward, while the useful reward appears when the agent enters the goal. A learning method must therefore use experience to improve actions that may occur several steps before the rewarding transition.

One Transition Two Uses

A real experience becomes reusable

Suppose the agent experiences one transition: it takes action a in state s, receives reward r, and arrives in state s'. How can this one event contribute to Dyna-Q twice?

Direct learning: The real transition is used immediately by the direct reinforcement-learning part of Dyna-Q, described in the source as one-step tabular Q-learning.

Model update: The experience also teaches the table-based model the recorded reward and next-state prediction associated with the experienced state-action pair.

Search control: During a later planning phase, search control selects a starting state and action from previously experienced state-action pairs.

Simulated experience: The model returns its recorded reward and next-state prediction for that selected pair, creating a simulated experience.

Planning update: The simulated experience is used by one-step tabular Q-planning, so the original real transition can influence learning again without requiring the agent to immediately repeat the environmental interaction.

A real transition supports an immediate direct-learning update and can later support one or more planning updates through the model.

real experiencerecord transitionavailable known pairsselect pairreturn predictionplanning updateEnvironmentreal transitionDirect learningone-step Q-learningModelrecorded reward and nextstateSearch controlknown state-action pairSimulated experiencemodel predictionPlanningone-step Q-planning
What happens after the agent experiences one transition, and how can that transition be reused?

The model does not replace the environment. It preserves information from real interaction so that the planning component can reuse that information later.

The Planning Budget n

In Dyna-Q, n specifies the number of planning steps performed per real step. It is useful to interpret n as a planning budget: n = 0 means no additional planning steps, while larger values add more simulated updates after each real interaction.

AgentPlanning after each real stepReported improvement
n = 0No planning; direct reinforcement learning onlyAbout 25 episodes
n = 5Five planning stepsAbout 5 episodes
n = 50Fifty planning stepsAbout 3 episodes
more planningmore planningn = 00 planning steps; about 25episodesn = 55 planning steps; about 5episodesn = 5050 planning steps; about 3episodes
How does increasing n change the number of simulated updates and the reported learning speed?

The comparison used the same kind of maze and varied how much planning occurred after each real step. The first episode was the same for every n value because the random-number-generator seed was held constant across the algorithms. That episode took about 1700 steps and was not shown in the reported curves.

How Planning Spreads Policy Information

Direct learning and planning differ in how quickly useful information can spread through the maze. With n = 0, the agent relies only on direct reinforcement learning. In the halfway-through-the-second-episode comparison, each episode adds only one more step to the learned policy, so only the final step has been learned at that point. With n = 50, the first episode still directly teaches only one step, but planning during the second episode develops a much larger policy while the agent is still wandering near the start. By the end of the second episode, the policy reaches almost back to the start state. By the end of the third episode, the source reports a complete optimal policy and perfect performance.

policy informationpolicy informationpolicy informationGoal reward+1 on entryNear-goal policyuseful information spreadsMiddle policyuseful information spreadsStart statereached sooner withplanning
How does planning propagate information from the goal toward the start more quickly than direct learning alone?

Planning spreads information because the model allows previously recorded transitions to be replayed as simulated experiences. Once useful information has been obtained near the goal, repeated planning updates can affect earlier parts of the learned policy while the agent is still exploring other parts of the maze.

Known Pairs and Search Control

Search control selects the starting states and actions for simulated experiences. Dyna-Q restricts these selections to state-action pairs that the agent has already experienced. That restriction matters because the table-based model has recorded reward and next-state information for experienced pairs. When search control selects one of those pairs, the model can return a simulated experience based on what was recorded.

lookup succeedsreturn predictionno recorded pairKnown state-actionpairselected by search controlModel lookuprecorded reward and nextstateSimulated experienceused for planningUnexperienced pairno recorded modeltransition
Why does planning begin with a previously experienced state-action pair?

Mistakes About Dyna-Q

  • Treating planning as a replacement for real interaction

    The comparison describes planning as additional computation performed after real steps. The agent still interacts with the maze and builds its model from real experience.

    Fix: Think of n as an additional planning budget applied to what the agent has already experienced.

  • Assuming n changes the maze

    The experiment compares different amounts of planning on the same kind of task.

    Fix: n changes the number of planning steps performed after each real step, not the maze structure.

  • Assuming the first episode proves that planning is faster

    The first episode was the same for every n value because the random-number-generator seed was held constant across the algorithms.

    Fix: Compare what happens after the shared initial experience, when the agents use different planning budgets.

  • Planning from arbitrary unexperienced pairs

    The table-based model has no recorded reward and next-state prediction for that pair.

    Fix: Let search control select from previously experienced state-action pairs.

  • Treating the reported n = 50 result as a universal rule

    The reported learning speeds came from a particular maze experiment and particular parameter settings.

    Fix: Use the result to understand the effect of additional planning in this experiment, not as a universal guarantee.

Check Your Understanding

MEDIUM

A Dyna-Q agent has just experienced a state-action-reward-next-state transition. Explain two different ways this transition can affect the agent: one immediate use and one later use. Then explain why a planning step should select a previously experienced state-action pair.

Hints
  • Separate direct reinforcement learning from model-based planning.
  • Identify what the model records from the real transition.
  • Ask whether the model has a prediction for an unexperienced pair.
MEDIUM

Compare n = 0 and n = 50 in the reported Dyna Maze experiment. Describe what happens after each real step and why the n = 50 agent can develop a larger policy while it is still wandering near the start.

Hints
  • n = 0 performs no planning steps.
  • n = 50 performs fifty planning steps after each real step.
  • Planning reuses transitions recorded in the model.

Key Takeaways

  1. The Dyna Maze has 47 states, four directional actions, zero reward on ordinary transitions, and reward +1 when entering the goal.
  2. Dyna-Q combines one-step tabular Q-learning for direct reinforcement learning with random-sample one-step tabular Q-planning.
  3. The parameter n is the number of planning steps performed per real step; n = 0 is the nonplanning baseline.
  4. A real transition can update learning directly and also teach a model that later produces simulated experiences.
  5. Planning can spread useful policy information through the maze faster than direct learning alone, as shown by the reported improvement for larger n values.

Key Takeaways

  • Dyna-Q learns from real interaction and also plans from a model built from that interaction.
  • The planning budget n controls how many simulated planning updates occur after each real step.
  • In the reported Dyna Maze experiment, larger planning values led to faster improvement: about 25 episodes for n = 0, about 5 for n = 5, and about 3 for n = 50.
  • Planning reuses experienced state-action pairs so that useful information can spread through the learned policy while the agent continues interacting with the maze.