State-Space Planning and Random-Sample Q-Planning
Random-sample Q-planning begins by selecting S and A at random.
Planning Through States
State-space planning treats each possible situation as a location in a search space. An agent considers how actions move it from one state to another, then uses value information to guide a policy or path toward a goal. The central planning question is therefore two-part: where will an action take the agent, and how valuable is the resulting state?
Random-sample Q-planning is a particular short planning process. It first selects a state S and an action A at random. It sends that pair to a sample model, which supplies a sampled reward R and a sampled next state S′. It then applies one-step tabular Q-learning to update the selected Q(S, A) entry. One planning step therefore turns one randomly selected state-action pair into one Q-value update.
Tracing One Planning Step
Consider a generated state set containing Entrance, Hallway, and Goal. These names are only an illustrative practice setting. Suppose a planning step randomly selects Hallway as S and an available action as A. At this moment, the selected pair identifies the Q(Hallway, A) entry that the step will address. The selection itself does not yet provide a reward or a next state.
The selected S and A are inputs to the sample model. The model returns the missing transition information: a sampled reward R and a sampled next state S′. This separation is important. The random selection identifies which Q-value entry is being addressed, while the model supplies the outcome information used by the update.
| Item | Role in the planning step |
|---|---|
| S | The randomly selected current state |
| A | The randomly selected action available in S |
| Q(S, A) | The selected Q-value entry addressed by the step |
| R | The sampled reward returned by the sample model |
| S′ | The sampled next state returned by the sample model |
The selected pair and the sampled transition have different roles.
Updating the Selected Entry
Before the update, the selected Q(S, A) entry has its current value. After the sample model returns R and S′, the one-step tabular Q-learning rule combines the current state, selected action, sampled reward, and sampled next state to produce a new value for that same selected entry. The operation changes Q(S, A); it does not turn the randomly selected pair into a different pair.
Tracing Entrance to Hallway
Trace one generated random-sample Q-planning step in which the randomly selected state is Entrance and the selected action leads, according to the sample model, to Hallway.
Select: Choose Entrance as S and choose an available action as A. This identifies Q(Entrance, A) as the entry to update.
Sample: Send S and A to the sample model. The model returns a sampled reward R and Hallway as the sampled next state S′.
Back up: Apply one-step tabular Q-learning to the selected Q(Entrance, A) entry using the current Q-value together with R and S′.
Distinguish: The value before the operation is the current Q(Entrance, A). The value after the operation is the updated Q(Entrance, A). They refer to the same selected entry at different points in the process.
One sampled transition has produced one update to the Q-value for the selected state-action pair.
State-Space Structure
State-space planning searches through states rather than through complete plans. Actions move the process from state to state, and value functions evaluate states. Those value functions act as an intermediate guide toward improving the policy: simulated experience is processed through backup operations, the resulting value information is used to improve the policy, and the policy guides choices among actions.
The shared structure is more important than the names of the generated states. A planning method uses simulated experience, applies backup operations to compute or update value functions, and uses those values as an intermediate step toward improving the policy. This is the common architecture of state-space planning methods described here.
Two Planning Spaces
| Feature | State-space planning | Plan-space planning |
|---|---|---|
| What is searched? | States | Plans |
| What does the search seek? | An optimal policy or path to a goal | A transformation from one plan to another |
| What moves the process? | Actions move from state to state | Operators transform one plan into another |
| Where are value functions defined? | Over states | Over plans |
| Examples named in the source | State-space planning methods | Evolutionary methods and partial-order planning |
Do not treat these labels as interchangeable. State-space planning searches states for an optimal policy or path to a goal. Plan-space planning searches a space of plans; its operators transform one plan into another, and its value functions are defined over plans rather than states. Partial-order planning is a plan-space example in which the ordering of steps can remain partly undetermined during some stages.
Common Misunderstandings
Treating the random state-action selection as if it already contained the reward and next state.
The selected S and A identify the Q-value entry, but the sample model supplies R and S′.
Fix:
Trace the process in order: select S and A, send them to the sample model, receive R and S′, then update Q(S, A).Confusing the current Q-value with the updated Q-value.
The current Q(S, A) is the value before the one-step operation; the updated Q(S, A) is the value after the operation.
Fix:
Use the words current and updated to mark the two different points in the trace.Assuming that the update changes every Q-value entry.
The described step addresses the selected Q(S, A) entry.
Fix:
Follow the randomly selected state and action and identify that single entry as the target of the step.Treating simulated experience as unrelated to value functions.
Simulated experience is processed through backup operations to compute value functions, which then help improve the policy.
Fix:
Place the backup between simulated experience and policy improvement in your mental model.Calling state-space planning plan-space planning.
State-space planning searches states and uses actions to move between them. Transforming plans belongs to plan-space planning.
Fix:
Ask what object is being searched: states in state-space planning, plans in plan-space planning.
Practice Trace
A random-sample Q-planning step selects Entrance as S and an available action as A. The sample model returns a reward R and Hallway as S′. Describe the role of each item and state which Q-value entry receives the one-step tabular Q-learning update.
Hints
- Separate the randomly selected inputs from the model's returned outputs.
- The selected state and action identify the Q-value entry.
- The reward and next state complete the transition information used by the update.
Practice Answer
Identify the roles of Entrance, A, R, Hallway, and the target Q-value.
Current state: Entrance is S, the randomly selected current state.
Selected action: A is the randomly selected action available in Entrance.
Model outputs: R is the sampled reward and Hallway is the sampled next state S′ returned by the sample model.
Update target: The one-step tabular Q-learning operation updates Q(Entrance, A), using the current entry together with the sampled transition information.
The sample model does not choose the Q-value entry. The random selection chooses Q(Entrance, A), and the model supplies the reward and next state used to update it.
Key Takeaways
- Random-sample Q-planning begins by randomly selecting a state S and an available action A.
- The selected pair identifies the Q(S, A) entry addressed by the planning step.
- A sample model receives S and A and returns a sampled reward R and next state S′.
- One-step tabular Q-learning uses the current Q-value and sampled transition information to produce an updated value for the selected entry.
- State-space planning searches states, uses actions to move between states, and uses value functions to help improve a policy; plan-space planning searches and transforms plans.
Key Takeaways
- Random-sample Q-planning separates selection from simulation: S and A are selected first, while R and S′ come from the sample model.
- The planning step updates the selected Q(S, A) entry using one-step tabular Q-learning.
- The current Q-value is the value before the operation; the updated Q-value is the result after the sampled transition is backed up.
- State-space planning uses simulated experience and backup operations to compute value functions that help improve a policy.
- State-space planning searches states, whereas plan-space planning searches plans.