Concepts / Predecessor States in Planning

Predecessor States in Planning

Backward focusing directs planning from a changed value toward states that may depend on it.

  • Programming

When One Value Changes

Planning does not always need to recompute everything after new information arrives. When the estimated value of one state changes, the useful question is: which other states and actions could depend on that change? Backward focusing answers this question by starting at the changed state and looking backward toward states that can lead into it.

Backward focusing directs planning from a changed value toward states that may depend on it.

value may propagatedirect incoming actionPredecessor 2may depend on P1Predecessor 1action leads to CChanged statevalue changed
How does a change in one state's value move backward through states that can transition into it?

The Backward Direction

A predecessor state is a state from which an action can lead into the state whose value changed. Backward focusing begins with that changed state, inspects the actions that lead directly into it, and then considers predecessor states when their values change. The direction is backward relative to the transition structure: instead of immediately examining every possible action, planning follows incoming connections to the changed state.

Backward focusing is a planning strategy that directs computation from a changed state value toward states and actions that may depend on that value.

startsstartsleads intoleads intoState AAction 1from State AState Cchanged valueState BAction 2from State B
Which states are predecessors of the changed state, and which actions connect them to it?

Selecting Immediate Backups

A Changed State with Three Candidate Actions

Suppose the value of State C has changed. Action 1 leads from State A to State C, Action 2 leads from State B to State C, and Action 3 is unrelated to State C.

Start at State C: Treat State C as the point from which backward focusing begins because its value changed.

Inspect incoming actions: Action 1 and Action 2 lead directly into State C, so they are the direct incoming actions.

Choose the immediate backups: The one-step backups for Action 1 and Action 2 are immediately useful. They are the backups whose action outcomes include the changed state.

Defer the unrelated action: Action 3 does not lead directly into State C, so it is not the immediate focus of this propagation step.

Continue if needed: If updating the direct incoming actions changes the values of State A or State B, those predecessor states can become the next sources of backward propagation.

Immediately useful backups are the backups for Action 1 and Action 2, followed by possible backups for predecessor states whose values change.

begincheck each actionyesnoif predecessor value changesChanged statevalue changedInspect incomingactionsLeads directly intochanged state?Use one-step backupimmediate focusInspect changedpredecessorscontinue backwardUnrelated actionnot immediate focus
Which predecessor-state backups become immediately relevant after a state value changes?

The immediate selection rule is local and directional: after a state changes in value, first use the backups for actions that lead directly into that state. Actions unrelated to the changed state are not the immediate focus of this propagation step.

Propagation Through Predecessors

Backward focusing is not limited to one backward step. Once the direct incoming actions are updated, their predecessor states become possible next sources of change. If a predecessor state's value changes as a result, planning can inspect the actions that lead into that predecessor. In this way, a value change can propagate backward through a chain of predecessor states.

Following a Chain Backward

Consider a chain in which State P2 has an action leading to State P1, and State P1 has an action leading to changed State C.

Changed state: Begin with State C because its value has changed.

First predecessor layer: The action from State P1 to State C is a direct incoming action, so its one-step backup is immediately useful.

Possible new change: If updating that backup changes the value associated with State P1, State P1 becomes the next changed predecessor state.

Second predecessor layer: Now inspect actions leading directly into State P1. The action from State P2 to State P1 is the next relevant incoming action.

The computation moves from C to P1 and, if needed, from P1 to P2. Each step follows direct incoming actions from the most recently changed state.

At each propagation step, reapply the same rule: identify the state whose value changed, find actions that lead directly into it, and focus on those one-step backups before moving farther backward.

Rewards Beyond Explicit Goals

SituationStarting point for backward focusingWhat matters next
Reward associated with an explicit goal stateThe goal state's changed valueActions that lead directly into that state
Reward associated with an ordinary stateThe ordinary state's changed valueActions that lead directly into that state

The method is general because the changed state does not have to be an explicit goal. If new reward information changes the value of an ordinary state, backward focusing still starts at that state, examines actions that lead directly into it, and follows predecessor states when their values change. The selection rule remains the same whether or not the state is called a goal.

Common Reasoning Mistakes

  • Recomputing every action immediately after one state changes.

    Backward focusing directs attention first to actions that lead directly into the changed state. Unrelated actions are not the immediate focus of that propagation step.

    Fix: Inspect incoming actions first, then continue toward predecessor states only when their values become relevant through the propagation.

  • Looking at actions that leave the changed state instead of actions that enter it.

    The immediate one-step backups are for actions that lead directly into the changed state.

    Fix: Reverse your attention: identify the states and actions that can reach the changed state.

  • Stopping after the first predecessor layer.

    Value changes can propagate backward through predecessor states when their values change.

    Fix: After each relevant update, check whether a predecessor state's value changed and, if so, inspect actions leading into that predecessor.

  • Assuming the changed state must be a goal.

    The method is general and does not require the changed state to be a goal.

    Fix: Use the changed value as the trigger, whether the state is an explicit goal or an ordinary state.

Practice the Selection Rule

EASY

A state named R has just changed in value. Action A leads directly into R. Action B leads into another state, and that other state is not known to have changed. Action C also leads directly into R. Which actions should be the immediate focus of backward focusing, and what should you inspect next if the update changes the value of the state before Action A?

Hints
  • Start with the actions whose outcomes lead directly into R.
  • Do not select an action merely because it exists in the planning structure.
  • If the state before Action A changes, inspect actions that lead directly into that predecessor state.

What do you think happens?

Which actions are immediately useful after the value of R changes?

  • Action A only
  • Action B only
  • Actions A and C
  • Actions A, B, and C
Reveal answer

Answer: Actions A and C

They are the actions that lead directly into the changed state R. Action B is not the immediate focus because it leads into another state.

Key Takeaways

  1. Backward focusing starts from a state whose value changed and directs planning toward states that may depend on it.
  2. The immediately useful one-step backups are the backups for actions that lead directly into the changed state.
  3. If predecessor values change, the same process can continue backward through earlier predecessor states.
  4. The method does not require the changed state to be an explicit goal; it also applies when reward information changes the value of an ordinary state.
  5. Unrelated actions are not the immediate focus of the current propagation step.

Key Takeaways

  • Backward focusing begins at a changed state value.
  • Direct incoming actions receive the first useful backups.
  • Value changes can continue backward through predecessor states.
  • The approach applies to ordinary reward-associated states as well as explicit goals.