Concepts / Random-Sample Q-Planning

Random-Sample Q-Planning

Incremental planning breaks planning into small steps instead of requiring one uninterrupted calculation.

  • Programming

Planning in Small Steps

A planning process does not always need to finish one large calculation before anything else can happen. Incremental planning breaks planning into small steps. Each step performs only part of the planning work, so the process can pause, change direction, or share time with other activities.

This flexibility is the central reason to use incremental planning. A large uninterrupted calculation commits computation to a plan before the process can respond to a new need. With incremental planning, only a small amount of computation is tied to any one stage. If the process is interrupted or redirected, little computation is wasted.

requirescan respond after each stepLarge calculationone uninterrupted planCommitted computationmore tied to one directionSmall planningstepsone step at a timeRedirect or stoplittle computation wasted
Why can small planning steps stop or redirect with little wasted computation when an exact solution is too large to calculate?

Planning Beside Other Activities

Incremental planning is useful when planning is not the only activity taking place. Small planning steps can be interleaved with acting and with learning the model. The planner does not need to complete a single large calculation before an action occurs or before new model information is learned.

thencan be followed bythenPlanning stepsmall computationActingreal activityModel learningnew informationPlanning stepanother small computation
How can small planning steps be interleaved with real actions and learning new model information?

One Random-Sample Step

Random-sample Q-planning is an incremental planning procedure with three parts. First, select a state S and an action A at random. Next, send S and A to a sample model. The model supplies a sample reward R and a sample next state S′. Finally, apply one-step tabular Q-learning to update the selected Q(S, A) entry.

thensend S and Areturnsprovides inputsSelect Srandom stateSelect Arandom available actionSample modeluses S and AR and S′sampled transitionUpdate Q(S, A)one-step Q-learning
What happens as the planner selects a state and action, queries the model, and performs one Q-learning update?

The random selection identifies which Q-table entry the planning step addresses. The selected S and A do not yet provide a reward or next state. Those transition results come from the sample model.

From Model to Update

inputinputreturnsreturnsSselected stateSample modelreceives S and ARsample rewardAselected actionS′sample next state
How does the sample model supply the next reward and next state after a state-action pair is selected?

The update has two distinct sources of information. The randomly selected S and A identify the Q-value to update. The sample model contributes R and S′. One-step tabular Q-learning then combines the current Q(S, A) with the sampled transition to produce the value after the update.

Tracing One Planning Step

Trace a generated random-sample Q-planning step without assigning numerical values.

Select the pair: Choose a state S and an action A at random. This determines the Q(S, A) entry addressed by the planning step.

Query the model: Send S and A to the sample model. It returns a sample reward R and a sample next state S′.

Apply the update: Use the current Q(S, A), together with R and S′, in the one-step tabular Q-learning update.

Store the result: The selected Q(S, A) entry now has the value produced by that update.

One random-sample planning step turns a randomly selected state-action pair and one model-supplied transition into one Q-value update.

Before and After the Update

identifiesidentifiesone-step updateused withused withSselected stateQ(S, A)current valueRsample rewardQ(S, A)value after updateAselected actionS′sample next state
What is the difference between the Q-value before the sampled update and the Q-value after the update?

The current Q-value is the value stored for the selected pair before the update begins. The updated Q-value is the new value stored for that same Q(S, A) entry after the one-step rule has used the sampled reward and sampled next state.

Mistakes in the Planning Step

  • Starting with a reward and next state instead of selecting S and A.

    Random-sample Q-planning begins by selecting a state and an action. The sample model supplies R and S′ only after receiving S and A.

    Fix: Select S and A first, send them to the sample model, and then use the returned R and S′.

  • Updating an unspecified Q-value.

    The selected state-action pair determines which Q-table entry the planning step addresses.

    Fix: Track S and A from the random selection through the update.

  • Confusing the current value with the value after the update.

    The current value exists before the one-step rule is applied. The updated value is the result after R and S′ have been used.

    Fix: Name the two stages explicitly: current Q(S, A), then Q(S, A) after the update.

  • Assuming incremental planning must finish a complete solution immediately.

    Incremental planning is designed to perform small steps and remain interruptible.

    Fix: Evaluate the approach by its flexibility and compatibility with acting and model learning, especially when the problem is too large to solve exactly.

Practice the Trace

MEDIUM

Describe one complete random-sample Q-planning step in the correct order. Your description must include the randomly selected state S, the randomly selected action A, the sample model, the returned reward R, the returned next state S′, the current Q(S, A), and the Q(S, A) value after the one-step update.

Hints
  • Begin with the state-action selection.
  • Separate the model inputs S and A from the model outputs R and S′.
  • State clearly which Q-value exists before the update and which value exists after it.
  1. Select a state S at random.
  2. Select an action A at random from the actions available in S.
  3. Send S and A to the sample model.
  4. Receive the sample reward R and sample next state S′.
  5. Apply one-step tabular Q-learning to the selected Q(S, A) entry.
  6. Distinguish the entry's current value from its value after the update.

Key Takeaways

  1. Incremental planning performs planning in small steps rather than one uninterrupted calculation.
  2. Small steps make planning easier to interrupt or redirect with little wasted computation.
  3. Incremental planning can be interleaved with acting and learning the model, and it can be useful when a problem is too large to solve exactly.
  4. Random-sample Q-planning selects S and A at random, asks a sample model for R and S′, and updates Q(S, A).
  5. The current Q-value is the value before the update; the updated Q-value is the value after one-step tabular Q-learning uses the sampled transition.

Key Takeaways

  • Incremental planning divides planning into small, interruptible steps.
  • Its flexibility allows planning to share time with acting and model learning.
  • Random-sample Q-planning first selects S and A, then obtains R and S′ from a sample model.
  • One planning step applies one-step tabular Q-learning to the selected Q(S, A) entry.
  • The current Q-value and the value after the update are different stages of the same selected entry.