Concepts / State-Space Planning

State-Space Planning

Incremental planning breaks planning into small steps instead of requiring one uninterrupted calculation.

  • Programming

Planning Without One Long Calculation

A planning process does not always need to complete one large calculation before anything else can happen. Incremental planning divides planning into small steps. This makes the process easier to interrupt, redirect, or combine with other activities such as acting in an environment and learning the model.

beginrepeatcheck needstop or change directionkeep planningPlanning taskPlanning stepsmall computationPlanning stepsmall computationInterruptionRedirected planninglittle wasted computationFurther planning
How does planning proceed one small step at a time, and what happens when the process is interrupted or redirected before the full plan is complete?

The central tradeoff is that incremental planning may not produce a complete solution immediately, but it keeps computation flexible. If the process must stop or change direction, less computation is tied to one unfinished stage.

Why Small Steps Matter

A large, uninterrupted planning process commits computation to a plan before the process can respond to a need to stop or change direction. Incremental planning limits the amount of computation committed to any one stage. This is useful when the agent may need to act before planning is complete, when the model is still being learned, or when the planning problem is too large to solve exactly.

Planning styleHow computation is organizedResponse to interruption
Large uninterrupted calculationA plan receives one extended calculation before other activities can interveneMore computation may be tied to a plan that must be abandoned or changed
Incremental planningPlanning is divided into small stepsPlanning can be interrupted or redirected with little wasted computation

States, Actions, and Policy

State-space planning searches through states rather than through complete plans. In reinforcement learning terms, a state represents a possible situation in the search space. Actions move the process from one state to another, value functions evaluate states, and a policy can use this value information to choose actions.

State-space planning organizes decision making around two linked questions: where will an action take the agent, and how valuable is the resulting state? The state is the central object connecting these questions. Actions determine movement between states, while value functions provide evaluations that can guide improvement of the policy.

available choicemoves processis evaluatedsupports improvementguides choiceCurrent stateValue functionevaluates statesActionPolicyguides action choiceNext state
How do a state and available actions determine the next state, and how do value functions and policies guide the choice of action?

From Entrance to Goal

Consider an agent whose possible states include Entrance, Hallway, and Goal.

1. Represent situations as states: Entrance, Hallway, and Goal are treated as locations in a state-space search.

2. Consider an action: An action can move the process from one of these states to another.

3. Evaluate the resulting state: A value function supplies value information for the states reached through simulated transitions.

4. Improve the policy: The updated value information can contribute to a policy that makes better action choices.

The example shows how state, action, value function, and policy form one connected planning structure without requiring a particular numerical formula or reinforcement-learning algorithm.

Simulated Experience and Backups

State-space planning methods use simulated experience rather than relying only on a directly observed transition. A simulated transition describes movement from one state to another according to the planning process. A backup operation then processes that simulated experience to update value information associated with states.

simulate actionproducesis processed byupdatessupports improvement ofStateSimulated transitionstate to stateSimulated experienceBackup operationprocess informationValue functionupdated value informationPolicyimproved guidance
How does a simulated transition produce information that is backed up to update a value estimate or policy?

A backup operation is not a separate plan that replaces the policy. It is the operation that processes simulated experience to compute or update value information. Those values are an intermediate step toward improving the policy.

Planning Alongside Acting

Incremental planning is suitable for intermixing planning with acting and learning the model. The process can take a small planning step, act when needed, incorporate model learning, and then continue planning. It does not have to finish one large calculation before these other activities occur.

actexperiencelearn modelsupport simulationinform next actionEnvironmentAgentModelIncremental planning
How does information move in a cycle between acting in the environment, learning the model, and planning from simulated experience?

The value of this arrangement is flexibility. Planning can share time with activities that change what should be planned next. If the process is redirected, the small amount of computation associated with the current step limits wasted work.

Two Search Spaces

FeatureState-space planningPlan-space planning
Search objectStatesPlans
What changesActions move the process from state to stateOperators transform one plan into another
Value functionsDefined over statesDefined over plans
Planning aimSearches through states for an optimal policy or path to a goalSearches through a space of plans
ExamplesState-space planning methodsEvolutionary methods and partial-order planning

The difference is not merely vocabulary. State-space planning treats situations in the world as the objects being searched, with actions connecting those situations. Plan-space planning treats plans themselves as the objects being searched, with operators transforming one plan into another. Partial-order planning can leave the ordering of steps partly undetermined during some stages of planning.

Misunderstandings to Avoid

  • Assuming incremental planning must finish a complete plan before the agent can do anything else.

    Incremental planning is specifically organized into small steps so it can be interrupted, redirected, or intermixed with acting and model learning.

    Fix: Think of planning as a sequence of partial computations that can share time with other activities.

  • Treating simulated experience as a real environmental action.

    Simulated experience is generated for planning and is processed through backup operations.

    Fix: Separate the simulated transition used for planning from action taken in the environment.

  • Treating a backup operation as the final policy.

    Backup operations update value information, and value functions serve as an intermediate step toward improving the policy.

    Fix: Track the sequence: simulated experience, backup operation, updated value information, and policy improvement.

  • Confusing state-space planning with plan-space planning.

    The two approaches search different objects and define value functions over different objects.

    Fix: Ask whether the search is moving among states or transforming plans.

Check Your Understanding

MEDIUM

An agent has used a model to simulate a transition from Entrance to Hallway. It processes that simulated experience with a backup operation and then uses the resulting value information to improve its policy. Identify the role of each part: Entrance, the simulated transition, the backup operation, the value information, and the policy.

Hints
  • Start by identifying which item is a state.
  • Ask what connects one state to another.
  • Separate the operation that processes information from the value information produced by that operation.
  • The final item should guide action choices rather than merely evaluate a state.

What do you think happens?

If an incremental planning process is redirected before completing a large plan, should all of its planning computation necessarily be wasted?

  • Yes, because incomplete planning has no value
  • No, because small steps limit computation tied to the abandoned direction
  • Yes, because only a complete policy can be used
  • No, because planning never needs a model
Reveal answer

Answer: No, because small steps limit computation tied to the abandoned direction.

The main benefit of incremental planning is flexibility: it can be interrupted or redirected with little wasted computation.

The Planning Pattern

  1. Incremental planning divides planning into small steps instead of requiring one uninterrupted calculation.
  2. Small steps make planning interruptible and redirectable, with little wasted computation when the direction changes.
  3. State-space planning searches through states: actions move between states, value functions evaluate states, and policies use value information to guide action choice.
  4. Simulated experience is processed by backup operations to update value information, which supports policy improvement.
  5. State-space planning searches states, whereas plan-space planning searches plans and transforms one plan into another.

Key Takeaways

  • Incremental planning makes planning a sequence of small, interruptible computations.
  • Its flexibility allows planning to be intermixed with acting and learning the model, and makes it useful for problems too large to solve exactly.
  • State-space planning connects actions, states, value functions, and policies through a search over possible states.
  • Simulated experience is processed through backup operations to update value information and support policy improvement.
  • State-space planning searches states, while plan-space planning searches and transforms plans.