Sample Model
Planning uses model-generated experience, while learning uses experience from the real environment.
Planning Without a New Interaction
In reinforcement learning, an agent can improve its value estimates in two different ways. Learning uses experience collected from the real environment. Planning uses experience generated by a model of the environment. A sample model is the part of this process that produces one possible response to a chosen state and action: a reward and a next state.
Planning can improve value estimates without waiting for another real interaction with the environment. It asks the model to generate simulated experience instead.
The One-Step Planning Path
Random-sample one-step tabular Q-planning combines random selection with a one-step tabular Q-learning update. First, the method selects a state S and an available action A at random. It sends that pair to the sample model. The model returns a sample reward R and a sample next state S'. The four quantities S, A, R, and S' then go into one-step tabular Q-learning, which changes the Q-function entry Q(S, A).
What the Sample Model Returns
A model of the environment predicts what the environment may do after receiving a state and an action. The prediction concerns both parts of the response: the next state and the next reward.
A sample model returns one possible outcome for the supplied state and action. In the planning update, that outcome is represented by a sample reward R and a sample next state S'. The outcome is selected according to the environment's probabilities, but the model reports only the one result used for this iteration.
Distribution and Sample Reports
| Model type | Response to a state and action | Information exposed |
|---|---|---|
| Distribution model | All possible next states and rewards with their probabilities | Complete probability description |
| Sample model | One next-state and reward outcome | One outcome drawn according to the probabilities |
A distribution model lays out the complete set of possible outcomes and their probabilities. A sample model gives one outcome selected according to those probabilities. Both support planning because both let an agent anticipate what may happen after a state and action.
The Dozen-Dice Difference
Rolling a Dozen Dice
Consider a process that rolls a dozen dice and records the sum. What does each model type report?
Distribution model: It returns all possible sums together with the probability of each sum.
Sample model: It simulates the rolls and returns one sum selected according to the underlying probability distribution.
Compare the reports: The underlying dice process is unchanged. Only the model's report changes: one exposes the entire distribution, while the other exposes one sampled result.
The distribution model reports every possible sum and its probability; the sample model reports one sum.
The dozen-dice example is not a comparison of two different probability distributions. It is a comparison of two ways to report the same underlying process.
Coverage and Convergence
Random selection supports convergence only when it provides sufficient coverage over time. To converge to the optimal policy for the model, every state-action pair must be selected an infinite number of times, and the learning-rate parameter α must decrease appropriately over time. Repeated model-generated updates then give the Q-function opportunities to incorporate information about every state-action pair.
Assuming that one model call describes the whole environment.
A sample model returns one outcome, whereas a distribution model exposes the complete probability description.
Fix:
Interpret the sample as one model-generated transition used for one planning update.Treating the sample model as the component that performs the Q-function update.
The model generates the transition sample; one-step tabular Q-learning processes that sample and changes the Q-function.
Fix:
Keep the roles separate: request R and S' from the model, then apply the one-step Q-learning backup.Assuming random selection alone guarantees convergence.
Convergence requires every state-action pair to be selected an infinite number of times.
Fix:
Check both required conditions: sufficient repeated coverage and an appropriately decreasing α.
Trace and Check
Suppose random-sample one-step tabular Q-planning selects a state S and an available action A. The sample model returns a reward R and a next state S'. Identify which component creates the simulated transition, which component changes Q(S, A), and what additional conditions are required for convergence.
Hints
- Separate the model's response from the Q-learning backup.
- The convergence conditions concern how often state-action pairs are selected and how α changes over time.
What do you think happens?
A model receives a state and an action. Which result should you expect from a sample model?
Reveal answer
Answer: One possible reward and one possible next state
A sample model supplies one outcome drawn according to the environment's probabilities. A distribution model supplies the complete set of possible outcomes and their probabilities.
Essential Takeaways
- Planning uses model-generated experience, while learning uses experience from the real environment.
- Random-sample one-step tabular Q-planning selects S and A, asks the sample model for R and S', and applies a one-step Q-learning update to Q(S, A).
- A sample model returns one possible outcome; a distribution model returns all possible outcomes and their probabilities.
- Distribution models are more capable because they can generate samples, while sample models may be easier to construct.
- Convergence to the model's optimal policy requires every state-action pair to be selected infinitely often and α to decrease appropriately.
Key Takeaways
- A sample model creates simulated experience for planning by returning one reward and one next state after receiving a state and action.
- The model-generated transition is then processed by one-step tabular Q-learning to update Q(S, A).
- A distribution model exposes the full probability description, while a sample model exposes one sampled outcome.
- The dozen-dice example shows the difference between reporting all sums with probabilities and reporting one sum.
- Convergence requires infinite selection of every state-action pair and an appropriately decreasing α.