State-Space Planning
Incremental planning breaks planning into small steps instead of requiring one uninterrupted calculation.
Planning Without One Long Calculation
A planning process does not always need to complete one large calculation before anything else can happen. Incremental planning divides planning into small steps. This makes the process easier to interrupt, redirect, or combine with other activities such as acting in an environment and learning the model.
The central tradeoff is that incremental planning may not produce a complete solution immediately, but it keeps computation flexible. If the process must stop or change direction, less computation is tied to one unfinished stage.
Why Small Steps Matter
A large, uninterrupted planning process commits computation to a plan before the process can respond to a need to stop or change direction. Incremental planning limits the amount of computation committed to any one stage. This is useful when the agent may need to act before planning is complete, when the model is still being learned, or when the planning problem is too large to solve exactly.
| Planning style | How computation is organized | Response to interruption |
|---|---|---|
| Large uninterrupted calculation | A plan receives one extended calculation before other activities can intervene | More computation may be tied to a plan that must be abandoned or changed |
| Incremental planning | Planning is divided into small steps | Planning can be interrupted or redirected with little wasted computation |
States, Actions, and Policy
State-space planning searches through states rather than through complete plans. In reinforcement learning terms, a state represents a possible situation in the search space. Actions move the process from one state to another, value functions evaluate states, and a policy can use this value information to choose actions.
State-space planning organizes decision making around two linked questions: where will an action take the agent, and how valuable is the resulting state? The state is the central object connecting these questions. Actions determine movement between states, while value functions provide evaluations that can guide improvement of the policy.
From Entrance to Goal
Consider an agent whose possible states include Entrance, Hallway, and Goal.
1. Represent situations as states: Entrance, Hallway, and Goal are treated as locations in a state-space search.
2. Consider an action: An action can move the process from one of these states to another.
3. Evaluate the resulting state: A value function supplies value information for the states reached through simulated transitions.
4. Improve the policy: The updated value information can contribute to a policy that makes better action choices.
The example shows how state, action, value function, and policy form one connected planning structure without requiring a particular numerical formula or reinforcement-learning algorithm.
Simulated Experience and Backups
State-space planning methods use simulated experience rather than relying only on a directly observed transition. A simulated transition describes movement from one state to another according to the planning process. A backup operation then processes that simulated experience to update value information associated with states.
A backup operation is not a separate plan that replaces the policy. It is the operation that processes simulated experience to compute or update value information. Those values are an intermediate step toward improving the policy.
Planning Alongside Acting
Incremental planning is suitable for intermixing planning with acting and learning the model. The process can take a small planning step, act when needed, incorporate model learning, and then continue planning. It does not have to finish one large calculation before these other activities occur.
The value of this arrangement is flexibility. Planning can share time with activities that change what should be planned next. If the process is redirected, the small amount of computation associated with the current step limits wasted work.
Two Search Spaces
| Feature | State-space planning | Plan-space planning |
|---|---|---|
| Search object | States | Plans |
| What changes | Actions move the process from state to state | Operators transform one plan into another |
| Value functions | Defined over states | Defined over plans |
| Planning aim | Searches through states for an optimal policy or path to a goal | Searches through a space of plans |
| Examples | State-space planning methods | Evolutionary methods and partial-order planning |
The difference is not merely vocabulary. State-space planning treats situations in the world as the objects being searched, with actions connecting those situations. Plan-space planning treats plans themselves as the objects being searched, with operators transforming one plan into another. Partial-order planning can leave the ordering of steps partly undetermined during some stages of planning.
Misunderstandings to Avoid
Assuming incremental planning must finish a complete plan before the agent can do anything else.
Incremental planning is specifically organized into small steps so it can be interrupted, redirected, or intermixed with acting and model learning.
Fix:
Think of planning as a sequence of partial computations that can share time with other activities.Treating simulated experience as a real environmental action.
Simulated experience is generated for planning and is processed through backup operations.
Fix:
Separate the simulated transition used for planning from action taken in the environment.Treating a backup operation as the final policy.
Backup operations update value information, and value functions serve as an intermediate step toward improving the policy.
Fix:
Track the sequence: simulated experience, backup operation, updated value information, and policy improvement.Confusing state-space planning with plan-space planning.
The two approaches search different objects and define value functions over different objects.
Fix:
Ask whether the search is moving among states or transforming plans.
Check Your Understanding
An agent has used a model to simulate a transition from Entrance to Hallway. It processes that simulated experience with a backup operation and then uses the resulting value information to improve its policy. Identify the role of each part: Entrance, the simulated transition, the backup operation, the value information, and the policy.
Hints
- Start by identifying which item is a state.
- Ask what connects one state to another.
- Separate the operation that processes information from the value information produced by that operation.
- The final item should guide action choices rather than merely evaluate a state.
What do you think happens?
If an incremental planning process is redirected before completing a large plan, should all of its planning computation necessarily be wasted?
Reveal answer
Answer: No, because small steps limit computation tied to the abandoned direction.
The main benefit of incremental planning is flexibility: it can be interrupted or redirected with little wasted computation.
The Planning Pattern
- Incremental planning divides planning into small steps instead of requiring one uninterrupted calculation.
- Small steps make planning interruptible and redirectable, with little wasted computation when the direction changes.
- State-space planning searches through states: actions move between states, value functions evaluate states, and policies use value information to guide action choice.
- Simulated experience is processed by backup operations to update value information, which supports policy improvement.
- State-space planning searches states, whereas plan-space planning searches plans and transforms one plan into another.
Key Takeaways
- Incremental planning makes planning a sequence of small, interruptible computations.
- Its flexibility allows planning to be intermixed with acting and learning the model, and makes it useful for problems too large to solve exactly.
- State-space planning connects actions, states, value functions, and policies through a search over possible states.
- Simulated experience is processed through backup operations to update value information and support policy improvement.
- State-space planning searches states, while plan-space planning searches and transforms plans.