Concepts / Dyna-Q Agent

Dyna-Q Agent

A Dyna-Q agent can fail when the environment changes after it has learned a model.

  • Programming

When the Maze Changes

A Dyna-Q agent can learn a successful route and still fail to use a better route later. The reason is that the agent plans from its learned model. If the environment changes but the model does not include that change, planning continues from incomplete information.

route aroundopensagent continues withLeft routeestablished pathLeft routeestablished behaviorBarrierblocks shorter routeOpen passageshorter route availableRight pathnot a shortcutRight shortcutopened after 3000 steps
What changes in the maze when a shorter path opens after 3000 steps, and how does the agent's previously learned behavior relate to the new layout?

In the Shortcut Maze, the agent initially learns a route around the left side of a barrier. After 3000 steps, a shorter path opens on the right. The environment now offers a better route, but the regular Dyna-Q agent continues with the established left-side route.

The Model Falls Behind

A Dyna-Q agent uses both learning and planning. Information from interaction with the environment contributes to a learned model, and planning uses that model to make decisions. In the Shortcut Maze, the important failure is not simply one bad action. The model itself does not contain the newly opened shortcut. As a result, planning has no representation of that connection to reason about.

producesupdatessuppliessupportsEnvironmentchanged mazeExperienceobserved interactionsLearned modelshortcut missingPlanninguses stored informationLeft routeestablished behavior
How does information flow from real experience into the learned model, and why can planning continue using outdated information after the environment changes?

The sequence identifies the modeling problem: the environment changes, but the model has not incorporated the shortcut. Planning can reason from what the model contains, not from an unrepresented route. Therefore, additional planning from the same stale model does not by itself reveal the missing connection.

Planning Repeats the Old Route

What do you think happens?

After the right-side shortcut opens, what is most likely to happen if planning repeatedly uses a model that still contains only the old route?

  • Planning immediately discovers the shortcut
  • Planning continues to use the information supporting the old route
  • Planning stops because the environment changed
  • The model automatically adds every possible route
Reveal answer

Answer: Planning continues to use the information supporting the old route

The model does not contain the shortcut, so more planning from that same model does not reveal it. The agent continues with the established left-side route.

guidessupportsrepeatssupportsOld modelleft route representedPlanning updateuses model contentsLeft routeremains supportedPlanning updatesame missing connectionLeft routecontinues
What happens during repeated planning updates when the model still predicts the old route, and how does that keep the agent from switching to the new shortcut?

Planning is powerful only with respect to the information available in the model. If repeated planning uses the same stale model, it can continue to support the established route. More planning does not automatically create knowledge of a shortcut that has never entered the model.

Why Exploration Can Miss the Shortcut

An ε-greedy policy does include exploratory action selection, but exploration does not guarantee that the agent will discover the new shortcut. The issue is not that the agent has no exploration at all. The issue is that discovery requires the particular sequence of exploratory actions that takes the agent through the newly opened route.

An occasional exploratory action may fail to place the agent on the right sequence of actions. Without reaching and learning the new route, the model remains missing the shortcut, and planning continues to use stale information.

  • Assuming that any exploration will discover the shortcut.

    The source example identifies a required sequence of exploratory actions, not merely the presence of occasional exploration.

    Fix: Distinguish between having an exploratory policy and actually discovering the specific new connection.

  • Assuming that more planning must reveal the optimal path.

    Planning reasons from the model, and the model is missing the new connection.

    Fix: Ask first whether the model contains the route that planning is expected to use.

  • Explaining the failure as one bad action only.

    The deeper issue is that the model itself is missing the shortcut, so repeated planning also remains limited.

    Fix: Trace the environment, the learned model, and the planning process together.

Planning Versus Optimality

SituationWhat planning can useLikely consequence
Model contains the relevant routeInformation representing that routePlanning can reason from it
Model misses the newly opened shortcutStale information about the established routePlanning does not reveal the missing shortcut by itself

The misconception is that combining learning with planning means a Dyna-Q agent will always find the optimal path. The Shortcut Maze disproves that stronger claim. The agent initially learns the left-side route, the environment later opens a shorter right-side route, and the regular agent does not switch because its model does not contain the shortcut.

Check Your Understanding

MEDIUM

Explain the Shortcut Maze failure in three linked statements: first describe what changes in the environment after 3000 steps, then state what is missing from the agent's model, and finally explain why additional planning does not automatically make the agent use the shorter path.

Hints
  • Mention the shorter right-side path.
  • State that the learned model does not contain the shortcut.
  • Connect planning to the information represented in that model.

Tracing the Shortcut Maze

Why does the regular Dyna-Q agent continue using the left-side route after a shorter right-side path opens?

Initial learning: The agent initially learns a path around the left side of a barrier.

Environmental change: After 3000 steps, the maze opens a shorter path on the right.

Model check: The learned model does not contain the newly opened shortcut.

Planning consequence: Planning continues from stale information, so it does not by itself reveal the missing route.

Exploration check: ε-greedy exploration does not guarantee the particular sequence of exploratory actions needed to discover and learn the shortcut.

The agent continues with the established left-side route because the environment changed before the model incorporated the new right-side connection.

Key Takeaways

  1. A regular Dyna-Q agent can fail when the environment changes after it has learned a model.
  2. In the Shortcut Maze, a shorter right-side path opens after 3000 steps, but the agent continues with the established left-side route.
  3. The model does not contain the shortcut, so planning uses stale information.
  4. ε-greedy exploration does not guarantee the sequence of actions needed to discover a new route.
  5. Planning is not the same as finding the optimal path when the model is incomplete or outdated.

Key Takeaways

  • A Dyna-Q agent plans from its learned model, so an environmental change can make that model stale.
  • In the Shortcut Maze, the right-side shortcut opens after 3000 steps, but the agent keeps using the established left-side route.
  • Exploration is not enough by itself: ε-greedy behavior may fail to produce the exact action sequence needed to discover the shortcut.
  • Repeated planning cannot reveal a route that the model does not represent.
  • Planning supports decisions from current knowledge; it does not guarantee the optimal path in a changed environment.