Planning and Real-Time Action Selection
Goal-directed agents pursue explicit goals through sensing and action.
Why Acting Is an Ongoing Problem
A goal-directed agent does more than observe a situation. It pursues an explicit goal by sensing what is happening, selecting an action, and continuing to respond as the environment changes. Reinforcement learning studies this complete interaction: the agent acts in an environment, receives a new situation, and must decide what to do next.
The loop has two directions of interaction. The environment provides aspects of the situation for the agent to sense, while the agent sends actions that influence the environment. The agent is therefore participating in the problem rather than merely observing it. The goal supplies direction, sensing supplies information, and action selection determines the response.
Tracing One Decision
An Agent Responds to a Changing Situation
Imagine an agent whose explicit goal is to reach a desired location while operating in an uncertain environment.
Set direction: The goal gives the agent a basis for deciding which responses are useful.
Sense: The agent obtains information about the current situation from the environment.
Select: Using the goal and the available sensory information, the agent chooses an action.
Act: The chosen action influences the environment instead of leaving it unchanged.
Continue: The environment produces a new situation. The agent senses that situation and makes another decision rather than assuming the previous action settled the whole problem.
Goal-directed behavior is a continuing cycle of sensing, action selection, environmental change, and renewed sensing.
This trace shows why an action cannot always be treated as the final answer. The action changes the situation, and the resulting observation becomes relevant to the next decision. The agent must continue operating from what it can sense rather than treating the environment as fully known or completely predictable.
Planning Within the Whole Agent
Planning and action selection matter in reinforcement learning because of their roles in a larger agent. Reinforcement learning is not mainly about solving one detached task, such as prediction or planning in isolation. Its broader question is how a goal-seeking agent can sense its surroundings, choose actions, and continue operating when the environment is uncertain.
| View | Main focus | What it leaves out if treated alone |
|---|---|---|
| Isolated subproblem | A capability such as prediction or planning | The complete cycle of sensing, acting, and receiving environmental feedback |
| Whole-agent perspective | A goal-seeking agent interacting with an uncertain environment | Nothing essential about the interaction loop is separated from the agent's ongoing behavior |
Real-Time Choice and Future Consequences
Planning and real-time action selection are connected by the need to choose an immediate action while considering what may happen next. The environment may not be fully known, and the result of an action may not be completely predictable. Consequently, action selection is an ongoing control problem: use the available sensory information, choose an action, observe the resulting situation, and continue from there.
What do you think happens?
An agent selects an action in an uncertain environment. Should it assume that one fully predictable result will occur and stop sensing, or should it continue interacting and making decisions?
Reveal answer
Answer: Continue interacting and make decisions from the resulting situation
The environment may produce different possible results after an action. The agent must use the new situation it can sense and continue operating without treating the environment as fully known.
The branching of possible outcomes is a representation of uncertainty, not a claim about a particular number of outcomes. Its teaching purpose is to show that an action can lead to an environment the agent did not completely determine in advance. The agent therefore needs a continuing loop rather than a single one-time decision.
Agents Inside Larger Systems
The word complete does not mean that an agent must be an entire organism or an entire robot. A component inside a robot or another behaving system can still be treated as a complete agent for its own interaction problem. That component interacts directly with the rest of the larger system and indirectly with the larger system's environment.
Consider a component inside a larger behaving system. For its own interaction problem, that component can sense relevant information, select actions, and receive consequences from the rest of the system. It does not need to be the entire system to have a complete agent-environment interaction at its chosen level of analysis.
Mistakes About the Agent Perspective
Treating reinforcement learning as only a prediction or planning task.
Reinforcement learning asks how a goal-seeking agent operates through complete interaction, not merely how one isolated capability works.
Fix:
Place every capability in the larger loop of goal, sensing, action selection, environmental change, and renewed sensing.Assuming that the agent only observes the environment.
The agent sends actions that affect the environment, while the environment provides new information in return.
Fix:
Remember the two-way relationship: the environment supplies observations, and the agent supplies actions.Assuming that an action has one completely predictable result.
The environment may be uncertain, so the agent must continue operating from the situation that results.
Fix:
Sense again after action selection and use the new situation for the next decision.Assuming that only an entire robot or organism can count as an agent.
A component can be treated as a complete agent for its own interaction problem.
Fix:
Define the interaction problem and identify what the component senses, how it acts, and how it receives consequences.
Check Your Understanding
Explain, in your own words, why reinforcement learning studies more than an isolated planning problem. Include the roles of the goal, sensory observation, action selection, environmental uncertainty, and feedback. Then explain how a component inside a larger behaving system could still be treated as a complete agent for its own interaction problem.
Hints
- Start with the direction supplied by an explicit goal.
- Describe how sensing and action selection are connected.
- Explain why an action leads to another situation that must be sensed.
- Distinguish the component's own interaction problem from the larger system that contains it.
Key Takeaways
- A goal-directed agent pursues an explicit goal through sensing and action.
- The agent-environment relationship is a connected loop: the goal gives direction, sensing provides information, action selection chooses a response, and the action changes what the agent will sense next.
- Reinforcement learning focuses on complete, interactive agents operating under environmental uncertainty rather than on an isolated prediction or planning task.
- Action selection is ongoing because the agent must continue from the situation produced by its previous action.
- An agent can be a component inside a larger behaving system and still be complete for its own interaction problem.
Key Takeaways
- Goal-directed agents use sensing and action to pursue explicit goals.
- Reinforcement learning studies the complete interaction between an agent and an uncertain environment.
- Planning and action selection are valuable because of their roles in ongoing agent behavior.
- Actions influence the environment, which produces new situations for the agent to sense.
- A component within a larger behaving system can still be treated as a complete agent for its own interaction problem.