Concepts / One-Step Tabular Q-Learning

One-Step Tabular Q-Learning

Planning uses model-generated experience, while learning uses experience from the real environment.

  • Programming

A Backup Without New Interaction

One-step tabular Q-planning improves value estimates by using simulated experience from a model. It does not need to wait for another real interaction with the environment for every update. Instead, it randomly selects a state and an available action, asks a sample model for a reward and next state, and then applies one-step tabular Q-learning to change the corresponding Q-function entry.

The central separation is simple: the sample model creates the simulated transition, while one-step tabular Q-learning processes that transition and changes Q(S, A).

From Selection to Update

A single planning iteration follows a fixed path. First, the algorithm selects a state S and an available action A at random. Next, it sends that state-action pair to the sample model. The model returns a sample reward R and a sample next state S'. Finally, the four quantities S, A, R, and S' are given to one-step tabular Q-learning, which performs one Q-function backup and changes the entry Q(S, A).

choosesend S and Agenerateprovide R and S'changeState Sselected at randomAction Aavailable actionSample modelreceives S and AR and S'sample reward and nextstateQ-learning backupuses S, A, R, and S'Q(S, A)updated entry
How does one random-sample planning iteration move from a state-action choice to a changed Q-function entry?

Tracing a Simulated Transition

Follow one planning iteration after the algorithm randomly selects a state S and an available action A.

Selection: The algorithm chooses the state-action pair S and A from the tabular set.

Model request: The selected pair is sent to the sample model rather than directly to the real environment.

Simulated experience: The sample model supplies a sample reward R and a sample next state S'.

Backup: One-step tabular Q-learning uses S, A, R, and S' to perform one Q-function backup.

Result: The Q-function entry Q(S, A) is changed using the learning update.

The planning iteration has converted a randomly selected state-action pair into one Q-function update without requiring another real interaction.

What the Sample Model Supplies

The sample model is responsible for generating the simulated transition. It receives the selected state S and action A, then supplies the sample reward R and sample next state S'. These outputs are not a second planning-specific backup. They are the transition information that one-step tabular Q-learning uses for its ordinary one-step backup.

inputinputproducesproducesused withused withSselected stateSample modelgenerates a transitionsampleRsample rewardQ(S, A)updated by one-step backupAselected actionS'sample next state
How does the sample model turn a selected state-action pair into the transition information used by the Q-function update?

Planning and Real Experience

Planning and learning differ in where their experience comes from. Planning uses model-generated experience: the sample model supplies the reward and next state after a selected state-action pair. Learning uses experience from the real environment. In both cases, the resulting transition information can be used by a Q-learning update, but planning obtains its transition from the model rather than waiting for another real environmental interaction.

generatesprovidesaffectsaffectsSample modelplanning experienceR and S'simulated transitionReal environmentlearning experienceR and S'real transitionQ-functionupdated by Q-learning
What differs between planning with model-generated experience and learning from real-environment experience?
AspectPlanningLearning
Source of experienceModel-generated experienceExperience from the real environment
Transition providerSample modelReal environment
EffectThe simulated transition is processed by a Q-learning backupThe real transition is processed by learning

Coverage and Convergence

Random selection alone does not guarantee useful planning. For convergence to the model's optimal policy, every state-action pair must be selected an infinite number of times, and the learning-rate parameter α must decrease appropriately over time. The coverage condition ensures that no state-action pair is permanently neglected. The decreasing-α condition is the other requirement associated with one-step tabular Q-learning in this setting.

provides coveragecontrols updatesconverge towardState-actionselectionevery pair selectedinfinitely oftenQ-valuesrepeatedly updatedModel's optimalpolicyconvergence targetαdecreases appropriately
What conditions allow repeated random-sample updates to converge to the model's optimal policy?
  • Assuming that random selection automatically provides convergence.

    The convergence condition requires repeated selection of every state-action pair, not merely the use of randomness.

    Fix: Check that the selection process provides the required coverage over time.

  • Treating the sample model as the component that updates Q(S, A).

    The model generates the simulated reward and next state; one-step tabular Q-learning performs the backup.

    Fix: Describe the model as the transition generator and Q-learning as the updater.

  • Confusing simulated experience with real-environment experience.

    Planning uses model-generated experience, whereas learning uses experience from the real environment.

    Fix: Identify the source of R and S' before describing the update as planning or learning.

  • Ignoring the learning-rate condition.

    Convergence requires α to decrease appropriately over time.

    Fix: State both convergence requirements: infinite selection of every state-action pair and appropriately decreasing α.

Check the Flow

EASY

Describe one complete random-sample one-step tabular Q-planning iteration in the correct order. Your answer should identify what is selected first, what the sample model returns, which four quantities are passed to one-step tabular Q-learning, and which Q-function entry changes.

Hints
  • Begin with the randomly selected state and available action.
  • The sample model returns a reward and a next state.
  • The updated entry is written as Q(S, A).
MEDIUM

A planning method selects some state-action pairs repeatedly but never selects one particular pair. Does the stated convergence condition hold? Explain why or why not, and include the required condition on α.

Hints
  • Compare the selection pattern with the requirement for every state-action pair.
  • Remember that convergence has two stated conditions.

Essential Takeaways

  1. Random-sample one-step tabular Q-planning selects a state and an available action, then asks a sample model for a reward and next state.
  2. The sample model generates the simulated transition; one-step tabular Q-learning performs the Q-function backup.
  3. The update changes Q(S, A) using S, A, R, and S'.
  4. Planning uses model-generated experience, while learning uses experience from the real environment.
  5. Convergence to the model's optimal policy requires every state-action pair to be selected infinitely often and α to decrease appropriately over time.

Key Takeaways

  • A planning iteration begins by randomly selecting a state and an available action.
  • The sample model supplies the simulated reward and next state.
  • One-step tabular Q-learning uses those values to update Q(S, A).
  • Planning differs from learning because its experience comes from a model rather than the real environment.
  • Convergence requires infinite selection of every state-action pair and an appropriately decreasing α.