Values and Policies
Planning can occur as part of choosing an action rather than as a permanently stored calculation.
Planning for the Current Choice
Planning does not always produce a permanent calculation that is kept for every future decision. An agent can plan specifically for the state it currently occupies, use that planning to choose an action, and then discard the values and policy produced for that decision. The important question is not whether the result lasts forever. It is whether the result helps the agent choose well now and whether reusing it later would be valuable.
From Simulation to Policy
The internal planning path can be described as a sequence. It begins with simulated experience. That simulated experience is used in backups, which produce values. The values then contribute to a policy. Here, the policy is a basis for choosing among the actions available in the current state. The chain is therefore simulated experience, then backups, then values, then a policy, followed by action selection.
Tracing One Current Decision
An agent is in a current state with several available choices. It plans specifically for those choices and then selects one action.
Simulate: The agent begins with simulated experience related to the current state and its available choices.
Back up: The simulated experience is used in backups.
Form values: The backups produce values for the choices under consideration.
Form a policy: Those values contribute to a policy, which provides a basis for choosing among the current state's actions.
Choose: The agent uses the policy to select an action for the current state.
The planning result has served its immediate purpose: it supported action selection for the current state.
Immediate Use and Discarding
In one approach, planning concentrates on the state that needs an action now. The resulting values and policy are used immediately to choose an action. After that choice, the agent may discard them. These results are specific to the current state and its choices; they are not general results for every state.
What do you think happens?
An agent has many possible states and is unlikely to encounter its exact current state again soon. Is discarding the current values and policy necessarily evidence that planning was useless?
Reveal answer
Answer: No, because the planning may have been intended only to select the current action.
Planning can be useful even when its immediate values and policy have a short lifetime. The results may be discarded because they served the current action-selection decision and are unlikely to be reused soon.
Storing for Future Reuse
A second approach stores planning results so that the agent is farther along if it returns to the same state later. The decision to store is guided by the likely future importance of that state. If a later return is more relevant, storing the values and policy can make the earlier planning effort useful again.
| Choice | Immediate purpose | When it may fit |
|---|---|---|
| Discard the result | Use values and policy to choose the current action | The same state is unlikely to appear again soon |
| Store the result | Keep the agent farther along for a later return | The same state is likely to matter again later |
| Use a mixed approach | Focus on the current action while keeping possible future value | Immediate action selection and later reuse both matter |
Storage is not the only alternative to complete discarding. A mixed approach can focus planning on the current state while also storing results for potential future use. This separates two purposes: choosing an action now and preserving useful work in case the same state matters later.
Discarding or Storing
The central decision is about reuse. Discarding is reasonable when the planning result is specific to a state that the agent may not encounter again soon. Storing is beneficial when a later return to that same state is relevant enough that reusing the earlier result could save planning effort. Neither choice is automatically correct in every situation.
Assuming that useful planning must always be stored.
The planning may have been useful precisely because it supported the current action selection. Its short lifetime does not cancel that purpose.
Fix:
Judge the result by whether it helped choose the current action and whether future reuse is likely to matter.Treating values and policy as general results for every state.
The values and policy described here are specific to the current state and its choices.
Fix:
Keep the scope clear: the planning chain produces a basis for choosing among the current state's actions.Assuming that discarding and storing are the only possible overall strategies.
A mixed approach can focus planning on the current state while storing results for potential future use.
Fix:
Separate immediate action selection from the possible future value of reusing the result.Storing every result without considering future reuse.
The value of storage depends on how likely the same state is to matter again.
Fix:
Consider the likelihood of a later return before deciding whether storage is worthwhile.
Apply the Decision
An agent plans for its current state. The plan moves from simulated experience through backups and values to a policy, and the agent uses that policy to select an action. The same state is unlikely to appear again soon. Explain why discarding the values and policy can still be a reasonable result of useful planning. Then explain how the decision would change if the same state were likely to matter again later.
Hints
- Trace the planning chain before discussing storage.
- Separate the purpose of choosing an action now from the possibility of reusing a result later.
- Use the likelihood of encountering the same state again as the deciding circumstance.
Choosing a Storage Strategy
Compare two situations: in the first, the agent has many possible states and is unlikely to revisit the current state soon; in the second, a later return to the same state is relevant.
Identify the current purpose: In both situations, planning can support choosing an action for the current state.
Trace the result: The planning chain runs from simulated experience through backups and values to a policy for the current state's choices.
Evaluate reuse in the first situation: If the same state is unlikely to appear again soon, discarding the values and policy after action selection may be reasonable.
Evaluate reuse in the second situation: If the same state is likely to matter again later, storing the result may allow the earlier planning effort to be useful again.
Consider a mixed approach: When both immediate selection and possible future reuse matter, the agent can focus planning on the current state while storing results for potential future use.
The storage decision follows the expected value of reuse, not a rule that all planning results must be kept or discarded.
Key Takeaways
- Planning can be part of choosing an action for the current state without producing a permanently stored calculation.
- The planning chain progresses from simulated experience through backups and values to a policy.
- The resulting values and policy are specific to the current state and its choices.
- Discarding the results may be reasonable when the same state is unlikely to matter again soon.
- Storing the results may be beneficial when the same state is likely to matter again, and a mixed approach can support both immediate use and future reuse.
Key Takeaways
- Planning can support current action selection even when its immediate results are not stored permanently.
- The internal progression is simulated experience, backups, values, and then a policy for the current state's choices.
- Values and policies may be discarded after action selection when future reuse of the same state is unlikely.
- Storing planning results can make earlier effort useful again when the same state is likely to matter later.
- A mixed approach can use planning immediately while preserving results for possible future reuse.