Concepts / Planning and Acting Under Time Constraints

Planning and Acting Under Time Constraints

Planning and response time compete when planning is performed during action selection.

  • Programming

The Timing Problem

An action-selecting system faces two timing demands. It may need time to decide what to do, and it may also need to respond quickly after a situation appears. When planning happens during action selection, these demands compete: the system cannot finish choosing its action until the planning work on the decision path is complete.

requiresdelaysEncountered statePlanningmore decision timeSelected actionlater response
What happens to response latency when more time is spent planning before choosing an action?

More planning before action selection means more waiting before the response when planning is performed after the situation is encountered.

Planning on the Decision Path

Planning within action selection follows a direct sequence. A system encounters a state, plans for that state, and then uses the planning result to choose an action. Because planning contributes directly to the action chosen for the current state, the response must wait for planning to finish.

triggersinformsleads toEncountered statePlanon the decision pathSelect actionExecute action
How does planning fit into the sequence from receiving a situation to selecting and executing an action?

A Deliberate Move

A chess-playing program encounters a position and is allowed to spend time deciding its move.

Encounter: The program receives the current chess position as the state requiring a response.

Plan: The program can spend seconds or minutes considering the move and may plan dozens of moves ahead.

Select: The program selects an action after the planning work has contributed to the decision.

Timing: The response is delayed while the program plans, but that waiting is acceptable when a fast response is not required.

Planning during action selection fits this situation because the application allows time for deliberation before the action is selected.

Planning Before the Situation

Background planning changes when the expensive decision work occurs. Instead of waiting for a newly encountered state and then planning, the system computes a policy ahead of time. When a new state appears, the system applies that policy so action selection can happen rapidly.

computesis ready foris handled bysupportsBackground planningComputed policyready ahead of timeNewly encounteredstateApply policyRapid actionselection
How can planning happen in advance so that an action can be selected quickly when a new state appears?

A Policy Ready in Advance

A system must choose an action rapidly whenever a new state appears.

Before response time matters: The system performs planning in the background and computes a policy ahead of time.

New state: A newly encountered state appears and now requires an action.

Policy application: The system applies the already computed policy instead of beginning the expensive planning work at that moment.

Action selection: Because the policy is ready, the system can select an action rapidly.

Background planning moves planning away from the immediate response path, supporting low-latency action selection.

Background planning does not remove the need for planning. It changes the timing arrangement so that a policy is available before the newly encountered state requires a response.

Two Timing Arrangements

triggersinformsproducessupportsEncountered stateBackground planningbefore new statePlanningbefore selectionComputed policyreadySelected actionresponse waitsSelected actionrapid application
What is the difference between planning during action selection and planning that occurs beforehand?
Planning within action selectionBackground planning
Planning begins after the state is encountered.Planning occurs ahead of the newly encountered state.
Planning contributes directly to the current action choice.A policy is computed before the current response is needed.
The response waits for planning to finish.The policy can be applied for rapid action selection.
Fits situations where fast responses are not required.Supports situations where low response latency matters.

Neither arrangement is universally preferable. Planning during action selection gives the decision process time to plan when waiting is acceptable. Background planning is useful when the system must respond rapidly and can compute a policy before the newly encountered state requires action.

Deliberation and Latency

selectsallowsdelaysCurrent stateCurrent stateSelected actionquick decisionExtended planningmore deliberationSelected actionlater response
How does spending more time considering possible actions affect response speed and the possible quality of the chosen action?

When a system spends more time planning during action selection, it gives the decision process more time to plan, but the current response occurs later. The source material establishes this timing tradeoff; it does not make a universal claim that a longer planning period always produces a better action. The correct choice depends on whether the application can tolerate waiting or instead prioritizes low latency.

Check Your Reasoning

MEDIUM

A system must respond rapidly when a newly encountered state appears. It can either begin planning after the state appears or compute a policy in advance. Which arrangement better matches the response requirement, and why?

Hints
  • Ask whether planning is on the immediate decision path.
  • Look for the strategy that computes a policy before the newly encountered state requires a response.
  • Explain what happens when the policy is applied.

What do you think happens?

If planning begins only after a new state appears, will the response wait for that planning to finish?

  • Yes
  • No
Reveal answer

Answer: Yes

When planning is part of action selection, it lies on the decision path. The response must wait for the planning work to finish before the action can be selected.

To choose between the two strategies, first identify the timing requirement. If waiting is acceptable, planning can occur as part of action selection. If low latency matters, plan ahead so that a policy is ready to apply when a new state appears.

Common Timing Mistakes

  • Treating planning time and response time as independent when planning occurs during action selection.

    Planning on the decision path must finish before the action can be selected.

    Fix: Recognize that additional planning time increases waiting when the planning starts after the state is encountered.

  • Assuming background planning means that no planning is performed.

    Background planning still computes a policy; it performs that work before the newly encountered state requires a response.

    Fix: Describe background planning as moving planning earlier so the policy is ready for rapid application.

  • Using one strategy as the universal answer.

    Neither strategy is universally preferable. Their suitability depends on response-latency requirements.

    Fix: Choose planning during action selection when fast responses are not required, and choose background planning when low latency matters.

Key Takeaways

  1. Planning and response time compete when planning is performed during action selection.
  2. Planning within action selection is suitable when the application can tolerate waiting, such as a chess-playing program that spends time considering a move.
  3. Background planning computes a policy ahead of time and applies it when a newly encountered state appears.
  4. Planning during action selection and background planning differ mainly in when the planning work occurs.
  5. The right strategy depends on whether the application prioritizes time for deliberation or rapid response.

Key Takeaways

  • Planning on the immediate decision path delays action selection until planning finishes.
  • When fast responses are not required, the system can plan as part of choosing an action.
  • Background planning computes a policy before a newly encountered state requires a response.
  • Applying a ready policy supports rapid action selection when low latency matters.
  • Neither strategy is always best; the choice depends on the application's timing demands.