Random-Sample Q-Planning
Incremental planning breaks planning into small steps instead of requiring one uninterrupted calculation.
Planning in Small Steps
A planning process does not always need to finish one large calculation before anything else can happen. Incremental planning breaks planning into small steps. Each step performs only part of the planning work, so the process can pause, change direction, or share time with other activities.
This flexibility is the central reason to use incremental planning. A large uninterrupted calculation commits computation to a plan before the process can respond to a new need. With incremental planning, only a small amount of computation is tied to any one stage. If the process is interrupted or redirected, little computation is wasted.
Planning Beside Other Activities
Incremental planning is useful when planning is not the only activity taking place. Small planning steps can be interleaved with acting and with learning the model. The planner does not need to complete a single large calculation before an action occurs or before new model information is learned.
One Random-Sample Step
Random-sample Q-planning is an incremental planning procedure with three parts. First, select a state S and an action A at random. Next, send S and A to a sample model. The model supplies a sample reward R and a sample next state S′. Finally, apply one-step tabular Q-learning to update the selected Q(S, A) entry.
The random selection identifies which Q-table entry the planning step addresses. The selected S and A do not yet provide a reward or next state. Those transition results come from the sample model.
From Model to Update
The update has two distinct sources of information. The randomly selected S and A identify the Q-value to update. The sample model contributes R and S′. One-step tabular Q-learning then combines the current Q(S, A) with the sampled transition to produce the value after the update.
Tracing One Planning Step
Trace a generated random-sample Q-planning step without assigning numerical values.
Select the pair: Choose a state S and an action A at random. This determines the Q(S, A) entry addressed by the planning step.
Query the model: Send S and A to the sample model. It returns a sample reward R and a sample next state S′.
Apply the update: Use the current Q(S, A), together with R and S′, in the one-step tabular Q-learning update.
Store the result: The selected Q(S, A) entry now has the value produced by that update.
One random-sample planning step turns a randomly selected state-action pair and one model-supplied transition into one Q-value update.
Before and After the Update
The current Q-value is the value stored for the selected pair before the update begins. The updated Q-value is the new value stored for that same Q(S, A) entry after the one-step rule has used the sampled reward and sampled next state.
Mistakes in the Planning Step
Starting with a reward and next state instead of selecting S and A.
Random-sample Q-planning begins by selecting a state and an action. The sample model supplies R and S′ only after receiving S and A.
Fix:
Select S and A first, send them to the sample model, and then use the returned R and S′.Updating an unspecified Q-value.
The selected state-action pair determines which Q-table entry the planning step addresses.
Fix:
Track S and A from the random selection through the update.Confusing the current value with the value after the update.
The current value exists before the one-step rule is applied. The updated value is the result after R and S′ have been used.
Fix:
Name the two stages explicitly: current Q(S, A), then Q(S, A) after the update.Assuming incremental planning must finish a complete solution immediately.
Incremental planning is designed to perform small steps and remain interruptible.
Fix:
Evaluate the approach by its flexibility and compatibility with acting and model learning, especially when the problem is too large to solve exactly.
Practice the Trace
Describe one complete random-sample Q-planning step in the correct order. Your description must include the randomly selected state S, the randomly selected action A, the sample model, the returned reward R, the returned next state S′, the current Q(S, A), and the Q(S, A) value after the one-step update.
Hints
- Begin with the state-action selection.
- Separate the model inputs S and A from the model outputs R and S′.
- State clearly which Q-value exists before the update and which value exists after it.
- Select a state S at random.
- Select an action A at random from the actions available in S.
- Send S and A to the sample model.
- Receive the sample reward R and sample next state S′.
- Apply one-step tabular Q-learning to the selected Q(S, A) entry.
- Distinguish the entry's current value from its value after the update.
Key Takeaways
- Incremental planning performs planning in small steps rather than one uninterrupted calculation.
- Small steps make planning easier to interrupt or redirect with little wasted computation.
- Incremental planning can be interleaved with acting and learning the model, and it can be useful when a problem is too large to solve exactly.
- Random-sample Q-planning selects S and A at random, asks a sample model for R and S′, and updates Q(S, A).
- The current Q-value is the value before the update; the updated Q-value is the value after one-step tabular Q-learning uses the sampled transition.
Key Takeaways
- Incremental planning divides planning into small, interruptible steps.
- Its flexibility allows planning to share time with acting and model learning.
- Random-sample Q-planning first selects S and A, then obtains R and S′ from a sample model.
- One planning step applies one-step tabular Q-learning to the selected Q(S, A) entry.
- The current Q-value and the value after the update are different stages of the same selected entry.