Planning with Simulated Experience
Planning can occur as part of choosing an action rather than as a permanently stored calculation.
A Decision for the Present State
Planning does not always have to produce a permanently stored calculation. An agent can plan specifically for the state it occupies now, use the result to choose an action, and then discard the values and policy produced for that decision. The planning was still useful because its purpose was immediate action selection.
The important question is not only what planning produces, but also how long the result needs to remain useful.
From Simulation to Action
The internal planning path can be traced as a sequence. Planning begins with simulated experience. That experience is used in backups, which produce values. The values then contribute to a policy, which provides a basis for choosing among the actions available in the current state. The result is a selected action for the present decision.
A Planning Loop
Planning can be viewed as a loop around the current decision. The agent starts with a state that needs an action, generates simulated experience for the choices available there, uses backups to obtain values, and lets those values shape a policy. The policy then supports the current action selection. This description emphasizes progression: simulated experience is not the endpoint; it is an input to later stages of the decision.
A Short-Lived Result
Choosing Between Current Choices
An agent is in a current state with several available choices. It plans specifically for those choices rather than creating a permanently stored result for every state.
Simulate: The agent begins with simulated experience connected to the current state and its available choices.
Back up: The simulated experience is used in backups, producing values for the decision.
Form a policy: Those values contribute to a policy that provides a basis for choosing among the current choices.
Select: The agent uses the policy to select one action for the current state.
Release: After the action is selected, the values and policy may be discarded because they were produced for this decision.
The planning effort supports a real action even though its immediate values and policy do not remain as stored results.
This abstract example illustrates the scope of the result. The planning output is not useless simply because it has a short lifetime. It was aimed at selecting the action required now. Discarding the output means that the agent does not retain that particular calculation after the decision.
Discard or Retain
| Choice | What happens | When it can make sense |
|---|---|---|
| Discard | The agent uses the values and policy for the current action, then does not retain them. | The agent has many possible states and is unlikely to encounter this exact state again soon. |
| Store | The agent keeps planning results so they can support a later decision. | A later return to the same state is relevant, making reuse valuable. |
| Mixed approach | The agent focuses planning on the current state while storing results for possible future use. | Immediate action selection matters, but future reuse may also matter. |
Choosing a Retention Strategy
The deciding circumstance is how likely the same state is to matter again. When an application has many states and a return to the exact current state may be unlikely for a long time, concentrating on the current action and discarding the resulting values and policy can be reasonable. When a later return is more relevant, storing the results can make the earlier planning effort useful again. A mixed approach can focus planning on the current state while still storing results for potential future use.
Assuming that discarded planning results were useless
The planning was aimed at the current action selection, so its usefulness does not depend on permanent storage.
Fix:
Judge the result by whether it supported the decision it was created for.Treating a current-state policy as a general policy for every state
The planning results are specific to the current state and its choices.
Fix:
Keep the scope of the values and policy tied to the state and choices that produced them.Storing every result without considering reuse
The benefit of storage depends on the likelihood that the same state will matter again.
Fix:
Consider future reuse before deciding whether to retain the result.
Check Your Reasoning
An agent plans for its current state, uses the resulting values and policy to choose an action, and then encounters a setting with many possible states where returning to this exact state soon is unlikely. Explain why discarding the values and policy can be reasonable. Then describe how your answer would change if returning to the same state later became relevant.
Hints
- Start with the purpose of the planning result.
- Trace the sequence from simulated experience to the selected action.
- Connect storage to the likelihood that the same state will matter again.
- A strong answer should explain that the planning result already supported the current action, so discarding it does not make the planning useless. If the same state is likely to matter again, storing the result can make the earlier planning effort useful for a later decision.
Key Takeaways
- Planning can be part of selecting an action for the current state rather than a permanently stored calculation.
- The planning chain progresses from simulated experience through backups and values to a policy.
- Values and policies produced for the current decision may be discarded after the action is selected.
- Discarding is more reasonable when the same state is unlikely to matter again soon.
- Storing results can be beneficial when a later return to the same state is relevant, and a mixed approach can support both immediate and future use.