Concepts / Sample Model

Sample Model

Planning uses model-generated experience, while learning uses experience from the real environment.

  • Programming

Planning Without a New Interaction

In reinforcement learning, an agent can improve its value estimates in two different ways. Learning uses experience collected from the real environment. Planning uses experience generated by a model of the environment. A sample model is the part of this process that produces one possible response to a chosen state and action: a reward and a next state.

Planning can improve value estimates without waiting for another real interaction with the environment. It asks the model to generate simulated experience instead.

usesinteracts withPlanningmodel-generated experienceSample modelsimulated responseLearningreal-environment experienceEnvironmentcollected response
What is the difference between experience generated by a model for planning and experience collected from the real environment for learning?

The One-Step Planning Path

Random-sample one-step tabular Q-planning combines random selection with a one-step tabular Q-learning update. First, the method selects a state S and an available action A at random. It sends that pair to the sample model. The model returns a sample reward R and a sample next state S'. The four quantities S, A, R, and S' then go into one-step tabular Q-learning, which changes the Q-function entry Q(S, A).

send S,Areturnsuse S,A,R,S'Select S and Arandom state-action pairSample modelreceives S and AR and S'sample reward and nextstateQ(S, A)one-step backup
How does the process move from a randomly selected state-action pair through the sample model to a Q-function update?

What the Sample Model Returns

A model of the environment predicts what the environment may do after receiving a state and an action. The prediction concerns both parts of the response: the next state and the next reward.

A sample model returns one possible outcome for the supplied state and action. In the planning update, that outcome is represented by a sample reward R and a sample next state S'. The outcome is selected according to the environment's probabilities, but the model reports only the one result used for this iteration.

inputinputreturnsreturnsState Scurrent situationAction Achosen actionSample modelselects one outcomeReward Rone possible rewardNext state S'one possible next state
After receiving a state and action, how does the sample model produce one possible reward and next state?

Distribution and Sample Reports

Model typeResponse to a state and actionInformation exposed
Distribution modelAll possible next states and rewards with their probabilitiesComplete probability description
Sample modelOne next-state and reward outcomeOne outcome drawn according to the probabilities

A distribution model lays out the complete set of possible outcomes and their probabilities. A sample model gives one outcome selected according to those probabilities. Both support planning because both let an agent anticipate what may happen after a state and action.

returnsreturnsDistribution modelall outcomes andprobabilitiesPossible outcomesprobability for eachSample modelone selected outcomeOne outcomedrawn according toprobabilities
What does each model return for a given state and action, and how does a probability distribution differ from one sampled outcome?

The Dozen-Dice Difference

Rolling a Dozen Dice

Consider a process that rolls a dozen dice and records the sum. What does each model type report?

Distribution model: It returns all possible sums together with the probability of each sum.

Sample model: It simulates the rolls and returns one sum selected according to the underlying probability distribution.

Compare the reports: The underlying dice process is unchanged. Only the model's report changes: one exposes the entire distribution, while the other exposes one sampled result.

The distribution model reports every possible sum and its probability; the sample model reports one sum.

distribution model reportssample model reportsDozen dicesum recordedAll sumsprobability for eachDozen dicesum recordedOne sumone sampled result
Given the same dozen-dice process, what output does the distribution model return compared with the sample model?

The dozen-dice example is not a comparison of two different probability distributions. It is a comparison of two ways to report the same underlying process.

Coverage and Convergence

Random selection supports convergence only when it provides sufficient coverage over time. To converge to the optimal policy for the model, every state-action pair must be selected an infinite number of times, and the learning-rate parameter α must decrease appropriately over time. Repeated model-generated updates then give the Q-function opportunities to incorporate information about every state-action pair.

  • Assuming that one model call describes the whole environment.

    A sample model returns one outcome, whereas a distribution model exposes the complete probability description.

    Fix: Interpret the sample as one model-generated transition used for one planning update.

  • Treating the sample model as the component that performs the Q-function update.

    The model generates the transition sample; one-step tabular Q-learning processes that sample and changes the Q-function.

    Fix: Keep the roles separate: request R and S' from the model, then apply the one-step Q-learning backup.

  • Assuming random selection alone guarantees convergence.

    Convergence requires every state-action pair to be selected an infinite number of times.

    Fix: Check both required conditions: sufficient repeated coverage and an appropriately decreasing α.

Trace and Check

MEDIUM

Suppose random-sample one-step tabular Q-planning selects a state S and an available action A. The sample model returns a reward R and a next state S'. Identify which component creates the simulated transition, which component changes Q(S, A), and what additional conditions are required for convergence.

Hints
  • Separate the model's response from the Q-learning backup.
  • The convergence conditions concern how often state-action pairs are selected and how α changes over time.

What do you think happens?

A model receives a state and an action. Which result should you expect from a sample model?

  • Every possible next state and reward with its probability
  • One possible reward and one possible next state
  • Only the current state
  • Only the chosen action
Reveal answer

Answer: One possible reward and one possible next state

A sample model supplies one outcome drawn according to the environment's probabilities. A distribution model supplies the complete set of possible outcomes and their probabilities.

Essential Takeaways

  1. Planning uses model-generated experience, while learning uses experience from the real environment.
  2. Random-sample one-step tabular Q-planning selects S and A, asks the sample model for R and S', and applies a one-step Q-learning update to Q(S, A).
  3. A sample model returns one possible outcome; a distribution model returns all possible outcomes and their probabilities.
  4. Distribution models are more capable because they can generate samples, while sample models may be easier to construct.
  5. Convergence to the model's optimal policy requires every state-action pair to be selected infinitely often and α to decrease appropriately.

Key Takeaways

  • A sample model creates simulated experience for planning by returning one reward and one next state after receiving a state and action.
  • The model-generated transition is then processed by one-step tabular Q-learning to update Q(S, A).
  • A distribution model exposes the full probability description, while a sample model exposes one sampled outcome.
  • The dozen-dice example shows the difference between reporting all sums with probabilities and reporting one sum.
  • Convergence requires infinite selection of every state-action pair and an appropriately decreasing α.