Concepts / Planning and Real-Time Action Selection

Planning and Real-Time Action Selection

Goal-directed agents pursue explicit goals through sensing and action.

  • Programming

Why Acting Is an Ongoing Problem

A goal-directed agent does more than observe a situation. It pursues an explicit goal by sensing what is happening, selecting an action, and continuing to respond as the environment changes. Reinforcement learning studies this complete interaction: the agent acts in an environment, receives a new situation, and must decide what to do next.

gives directionprovides informationinformschoosesinfluencesGoaldirection for behaviorSensinginformation about thesituationAction selectionresponse to the situationActioninfluence on theenvironmentEnvironmentnew situation
How do an agent's goals, sensory observations, action selection, and environmental feedback connect over time?

The loop has two directions of interaction. The environment provides aspects of the situation for the agent to sense, while the agent sends actions that influence the environment. The agent is therefore participating in the problem rather than merely observing it. The goal supplies direction, sensing supplies information, and action selection determines the response.

Tracing One Decision

An Agent Responds to a Changing Situation

Imagine an agent whose explicit goal is to reach a desired location while operating in an uncertain environment.

Set direction: The goal gives the agent a basis for deciding which responses are useful.

Sense: The agent obtains information about the current situation from the environment.

Select: Using the goal and the available sensory information, the agent chooses an action.

Act: The chosen action influences the environment instead of leaving it unchanged.

Continue: The environment produces a new situation. The agent senses that situation and makes another decision rather than assuming the previous action settled the whole problem.

Goal-directed behavior is a continuing cycle of sensing, action selection, environmental change, and renewed sensing.

informsinfluencesproducesinformsObservationcurrent situationActionselected responseEnvironmentaffected by the actionNew observationresulting situationRevised responsenext action selection
How can the environment change after an action, and how does the agent use the resulting observation to revise its behavior?

This trace shows why an action cannot always be treated as the final answer. The action changes the situation, and the resulting observation becomes relevant to the next decision. The agent must continue operating from what it can sense rather than treating the environment as fully known or completely predictable.

supportsmay producemay produceleads toleads toCurrent situationwhat the agent can senseSelected actionimmediate responseOutcome Apossible resultNext situationnew information forcontinued controlOutcome Bpossible result
What happens when an agent must choose an immediate action while predicting uncertain future outcomes?

Planning Within the Whole Agent

Planning and action selection matter in reinforcement learning because of their roles in a larger agent. Reinforcement learning is not mainly about solving one detached task, such as prediction or planning in isolation. Its broader question is how a goal-seeking agent can sense its surroundings, choose actions, and continue operating when the environment is uncertain.

includesincludesincludesincludesIsolated taskprediction or planningaloneGoalexplicit directionComplete agentgoal, sensing, action,feedbackSensinginformation fromenvironmentActioninfluence on environmentUncertaintyincomplete certainty
What is included when reinforcement learning considers the complete agent-environment interaction rather than an isolated subproblem?
ViewMain focusWhat it leaves out if treated alone
Isolated subproblemA capability such as prediction or planningThe complete cycle of sensing, acting, and receiving environmental feedback
Whole-agent perspectiveA goal-seeking agent interacting with an uncertain environmentNothing essential about the interaction loop is separated from the agent's ongoing behavior

Real-Time Choice and Future Consequences

Planning and real-time action selection are connected by the need to choose an immediate action while considering what may happen next. The environment may not be fully known, and the result of an action may not be completely predictable. Consequently, action selection is an ongoing control problem: use the available sensory information, choose an action, observe the resulting situation, and continue from there.

What do you think happens?

An agent selects an action in an uncertain environment. Should it assume that one fully predictable result will occur and stop sensing, or should it continue interacting and making decisions?

  • Assume one fully predictable result and stop sensing
  • Continue interacting and make decisions from the resulting situation
Reveal answer

Answer: Continue interacting and make decisions from the resulting situation

The environment may produce different possible results after an action. The agent must use the new situation it can sense and continue operating without treating the environment as fully known.

The branching of possible outcomes is a representation of uncertainty, not a claim about a particular number of outcomes. Its teaching purpose is to show that an action can lead to an environment the agent did not completely determine in advance. The agent therefore needs a continuing loop rather than a single one-time decision.

Agents Inside Larger Systems

The word complete does not mean that an agent must be an entire organism or an entire robot. A component inside a robot or another behaving system can still be treated as a complete agent for its own interaction problem. That component interacts directly with the rest of the larger system and indirectly with the larger system's environment.

containsusesperformsinteracts indirectly withLarger behavingsystemoverall behaviorAgent componentcomplete for itsinteraction problemSensinginformation from directinteractionAction selectionresponse within the systemSystem environmentindirectly encountered
How does an agent's sensing and action selection fit inside the behavior of a larger system?

Consider a component inside a larger behaving system. For its own interaction problem, that component can sense relevant information, select actions, and receive consequences from the rest of the system. It does not need to be the entire system to have a complete agent-environment interaction at its chosen level of analysis.

Mistakes About the Agent Perspective

  • Treating reinforcement learning as only a prediction or planning task.

    Reinforcement learning asks how a goal-seeking agent operates through complete interaction, not merely how one isolated capability works.

    Fix: Place every capability in the larger loop of goal, sensing, action selection, environmental change, and renewed sensing.

  • Assuming that the agent only observes the environment.

    The agent sends actions that affect the environment, while the environment provides new information in return.

    Fix: Remember the two-way relationship: the environment supplies observations, and the agent supplies actions.

  • Assuming that an action has one completely predictable result.

    The environment may be uncertain, so the agent must continue operating from the situation that results.

    Fix: Sense again after action selection and use the new situation for the next decision.

  • Assuming that only an entire robot or organism can count as an agent.

    A component can be treated as a complete agent for its own interaction problem.

    Fix: Define the interaction problem and identify what the component senses, how it acts, and how it receives consequences.

Check Your Understanding

MEDIUM

Explain, in your own words, why reinforcement learning studies more than an isolated planning problem. Include the roles of the goal, sensory observation, action selection, environmental uncertainty, and feedback. Then explain how a component inside a larger behaving system could still be treated as a complete agent for its own interaction problem.

Hints
  • Start with the direction supplied by an explicit goal.
  • Describe how sensing and action selection are connected.
  • Explain why an action leads to another situation that must be sensed.
  • Distinguish the component's own interaction problem from the larger system that contains it.

Key Takeaways

  1. A goal-directed agent pursues an explicit goal through sensing and action.
  2. The agent-environment relationship is a connected loop: the goal gives direction, sensing provides information, action selection chooses a response, and the action changes what the agent will sense next.
  3. Reinforcement learning focuses on complete, interactive agents operating under environmental uncertainty rather than on an isolated prediction or planning task.
  4. Action selection is ongoing because the agent must continue from the situation produced by its previous action.
  5. An agent can be a component inside a larger behaving system and still be complete for its own interaction problem.

Key Takeaways

  • Goal-directed agents use sensing and action to pursue explicit goals.
  • Reinforcement learning studies the complete interaction between an agent and an uncertain environment.
  • Planning and action selection are valuable because of their roles in ongoing agent behavior.
  • Actions influence the environment, which produces new situations for the agent to sense.
  • A component within a larger behaving system can still be treated as a complete agent for its own interaction problem.