Concepts / Model-Based Processes

Model-Based Processes

Dyna-Q combines one-step tabular Q-learning for direct reinforcement learning with random-sample one-step tabular Q-planning.

  • Programming

One Experience, Two Uses

Dyna-Q uses a real interaction with an environment in two connected ways. First, the observed transition contributes directly to reinforcement learning. Second, the same experience teaches a table-based model. Later, that model can produce a simulated experience, allowing another learning update without requiring an immediate new interaction with the environment.

updatesteachesstores candidatesselectsreturnsupdatesReal transitionstate, action, reward, nextstateDirect learningQ updateSearch controlexperienced state-actionpairSimulated experiencemodel responsePlanningQ updateModelrecorded reward and nextstate
How does one observed transition update Q directly, enter the model, and later generate a simulated Q-learning update?

The important connection is that direct learning and planning are not separate sources of information. Planning reuses information first obtained through real experience and stored by the model.

Tracing a Real Transition

From a Real Step to a Planned Step

Suppose an agent experiences the transition from state A after taking action left: it receives reward 4 and reaches state B. How can this one transition support both direct learning and later planning?

Observe: The agent obtains a real transition containing the starting state A, the selected action left, the observed reward 4, and the next state B.

Learn directly: The transition is used by one-step tabular Q-learning for direct reinforcement learning, so the Q values are updated from the real experience.

Teach the model: The table-based model records the reward and next-state prediction associated with the experienced state-action pair A and left.

Select for planning: Search control can later select the experienced pair A and left as the starting point for a simulated experience.

Generate simulated experience: The model returns its recorded reward 4 and next-state prediction B for A and left.

Plan: The model-generated transition is used by random-sample one-step tabular Q-planning, producing another Q-learning update without needing a new immediate interaction for this particular transition.

The real transition has two effects: it updates Q through direct reinforcement learning and supplies a model record that can later support planning.

producesupdatesteachesprovides experienced pairschoosesreturns predictionupdatesEnvironmentreal interactionTransitionstate, action, reward, nextstateDirect learningone-step Q-learningSearch controlstate-action selectionSimulated experiencerecorded reward and nextstatePlanningone-step Q-planningModeltable-based record
How does information flow from a real environment interaction through direct Q-learning and the model into simulated planning updates?

Four Connected Roles

Dyna-Q becomes easier to understand when its four roles are kept distinct. The model stores what the agent learned about experienced state-action pairs. Search control chooses which stored state-action pair will begin a simulated experience. The simulated experience is the model's returned reward and next-state prediction. Planning applies one-step tabular Q-planning to that simulated transition. Together, these roles let the agent continue reinforcement learning from a model built through direct interaction.

queriesgeneratessupplies transitionupdatesModelrecorded reward and nextstateSimulated experiencemodel responsePlanningone-step Q-planningQ valuesupdated estimatesSearch controlstarting state and action
What does each Dyna-Q component contain, and how are model lookup, state-action selection, simulated experience generation, and planning connected?
Direct reinforcement learningModel-based planning
Uses a transition produced by direct interaction with the environmentUses a transition returned by the model
Updates Q from real experienceUpdates Q from simulated experience
Also teaches the modelReuses what the model has recorded

Why Planning Uses Known Pairs

Search control restricts planning samples to state-action pairs that have already been experienced. The reason follows directly from the model's role: the model can return a reward and next-state prediction only for a state-action pair whose transition has been recorded. An unexperienced pair has no recorded model response in the information described by Dyna-Q, so it cannot provide the simulated transition required for planning.

checkyesreturnssuppliesnoState-action pairExperiencedrecorded in modelModel lookupreward and next stateSimulated experiencePlanningQ updateUnexperiencedno recorded model response
How does search control select only state-action pairs stored in the model, and why are unexperienced pairs excluded from planning?

A state-action pair can be part of the agent's possible behavior without being eligible for model-based planning yet. Eligibility depends on whether the pair has been experienced and therefore has a recorded reward and next-state prediction in the model.

maps tomaps toreturned withreturned withState-action pairexperienced pairRewardrecorded valueSimulated experiencereward and next stateNext staterecorded prediction
What does the model store for each experienced state-action pair, and how does that stored reward and next state become a simulated experience?

Predict the Planning Path

What do you think happens?

An agent has directly experienced the pair state C and action right. It observed reward 2 and next state D. Which sequence best describes a later planning use of this experience?

  • The model selects a new real action in the environment
  • Search control selects the experienced pair, the model returns reward 2 and next state D, and planning updates Q
  • Planning selects an unexperienced pair and invents its reward
  • The model changes the recorded transition before Q-learning uses it
Reveal answer

Answer: Search control selects the experienced pair, the model returns reward 2 and next state D, and planning updates Q

The model supplies simulated experience for an experienced state-action pair. That simulated transition is then used by one-step tabular Q-planning.

MEDIUM

Explain in your own words why a transition can be used twice in Dyna-Q without representing two separate real interactions.

Hints
  • Identify what happens immediately after the real transition.
  • Identify what the model records.
  • Distinguish a real transition from the simulated transition later returned by the model.

Mistakes to Avoid

  • Treating direct learning and planning as unrelated processes

    The model is taught by real experience, and planning reuses the model's recorded reward and next-state prediction.

    Fix: Trace the real transition along both paths: direct Q-learning and model construction.

  • Confusing the model with the simulated experience

    The model is the table-based source of recorded information; the simulated experience is what the model returns for a selected experienced pair.

    Fix: Separate the lookup process from the returned reward and next-state prediction.

  • Allowing any possible state-action pair into planning

    The model has no recorded reward and next-state prediction for that pair.

    Fix: Remember that search control selects from previously experienced state-action pairs.

  • Assuming planning requires another immediate environment interaction

    The model can later generate simulated experience from the stored transition.

    Fix: Use the model's returned reward and next state as the input to planning.

Dyna-Q in One Pass

  1. The agent interacts directly with the environment and observes a transition.
  2. One-step tabular Q-learning uses that real transition for direct reinforcement learning.
  3. The transition teaches the table-based model by associating an experienced state-action pair with its recorded reward and next-state prediction.
  4. Search control selects a previously experienced state-action pair for a simulated experience.
  5. The model returns the recorded reward and next-state prediction for that pair.
  6. Random-sample one-step tabular Q-planning uses the simulated transition for another Q-learning update.

Dyna-Q combines direct reinforcement learning with model-based planning. Its model is built from real experience, search control chooses which stored experience to revisit, and planning learns from the simulated transition returned by the model.

Key Takeaways

  • Dyna-Q combines one-step tabular Q-learning for direct reinforcement learning with random-sample one-step tabular Q-planning.
  • A real transition both updates Q directly and teaches the table-based model.
  • The model returns a recorded reward and next-state prediction for an experienced state-action pair.
  • Search control selects the starting state and action for a simulated experience.
  • Planning is restricted to previously experienced state-action pairs because only those pairs have recorded model responses.