Agent-Environment Boundary and Policy
Locate the boundary by asking what the agent can change arbitrarily, not by tracing the body's outer surface.
The Boundary Is Not the Shell
Imagine studying a robot that chooses actions to complete a task. A natural first guess is that the agent is everything inside the robot's shell. For an animal, the same guess would identify the agent with everything inside the animal's skin. Reinforcement learning uses a more useful question: what can the agent change arbitrarily? The answer to that question, rather than the body's outer surface, determines where the agent-environment boundary is drawn.
Tracing the Control Boundary
The first panel represents the tempting physical-boundary view: everything inside the body is called the agent. The second panel applies the control test. The decision-making rule is treated as the agent, while motors and sensing hardware are treated as environmental components when the agent cannot change them arbitrarily. The boundary is therefore a modeling choice tied to the task, not a line copied from the robot's shell.
Absolute Control as the Test
Use absolute control as the main classification test. If a component lies outside the agent's arbitrary control, treat it as part of the environment for the task under study. This can include motors, sensing hardware, muscles, and reward computation. These components may be physically attached to the robot or physically located inside an animal, but physical attachment does not give the agent arbitrary control over them.
Classifying a robot component
A decision-making agent selects actions for a robot. The robot contains a motor and sensing hardware. The agent can select actions but cannot alter the motor or sensing hardware arbitrarily. Where should these components be placed for this task?
Identify the task: The boundary is being chosen for the particular decision-making task represented by the agent's actions and rewards.
Apply the control test: The agent does not have arbitrary control over the motor or sensing hardware.
Classify the components: The motor and sensing hardware are treated as part of the environment for this task, even though they are physically inside the robot.
The agent-environment boundary is drawn according to arbitrary control, not according to the robot's physical outline.
Control Is Not Knowledge
The control boundary should not be confused with a knowledge boundary. The environment does not have to be mysterious or unknown to the agent. An agent may understand how the environment works, including how rewards are computed from actions and states, while still lacking the ability to change those mechanisms arbitrarily.
Suppose an agent understands exactly how a task's reward is computed from actions and states. That understanding does not make reward computation part of the agent. The agent may know the rule while still being unable to rewrite it at will. The task can remain difficult because the agent must still find actions that accomplish the task.
Reward Computation Outside the Agent
Reward computation belongs to the environment because it defines the task and is not under the agent's arbitrary control. The agent can know that the reward is computed from actions and states, and it can use that knowledge when selecting actions. Knowing the computation, however, does not give the agent the power to change the reward rules.
Several Boundaries in One Robot
A complicated robot does not always have one permanently fixed agent-environment boundary. The boundary is selected after choosing the states, actions, and rewards for the particular decision-making task. Different agents can therefore operate at different levels within the same robot. A boundary useful for one task may be located differently for another purpose.
The important point is not to search for the one physically correct boundary. Instead, specify the decision-making task first. Then identify the relevant states, actions, and rewards, and apply the absolute-control test to that task. The same physical component can be treated differently in two models when the tasks and levels of decision-making differ.
Policy Selects the Action
A policy is the agent's action-selection rule: it uses the information available for the decision-making task to select an action. The policy is associated with the agent, while the consequences of the action, the physical mechanisms that carry it out, and the reward computation can remain part of the environment when they are outside the agent's arbitrary control.
This distinction keeps the model precise. The policy determines what action the agent selects, but selecting an action is not the same as arbitrarily controlling every mechanism involved in carrying it out. The robot's motor, sensing hardware, muscles, or reward computation may remain environmental components in the chosen model.
Common Classification Mistakes
Defining the agent as everything inside the robot's shell or animal's skin.
The physical outline is not the rule used to locate the agent-environment boundary.
Fix:
Ask whether the agent can change the motor arbitrarily. If it cannot, treat the motor as environmental for the task.Assuming that a known mechanism must belong to the agent.
Knowledge and control are different. The agent can know the reward rule without being able to alter it arbitrarily.
Fix:
Use absolute control, not secrecy or uncertainty, as the boundary test.Assuming that one boundary must apply to every analysis of a complicated robot.
The useful boundary depends on the states, actions, rewards, and decision-making task under study.
Fix:
Choose the task first, then select the boundary that matches its control relationships.
When classifying a component, write two separate notes: what the agent knows about the component and what the agent can change arbitrarily. The first note describes knowledge; the second determines the boundary.
Boundary Classification Practice
For each item below, decide whether it should be treated as part of the agent or the environment for a specified decision-making task. Justify your decision using absolute control, not physical location or knowledge: a policy, a motor, sensing hardware, and reward computation.
Hints
- First identify which item selects the action.
- Then ask whether the agent can alter each other item arbitrarily.
- Remember that knowing how reward computation works does not imply control over it.
Checking your classification
An agent selects actions through a policy. The motor, sensing hardware, and reward computation are outside the agent's arbitrary control, although the agent may understand how each works. Classify the four items.
Policy: Treat the policy as part of the agent because it is the action-selection rule for the decision-making task.
Motor: Treat the motor as environmental when the agent cannot change it arbitrarily.
Sensing hardware: Treat sensing hardware as environmental when it is outside the agent's arbitrary control.
Reward computation: Treat reward computation as environmental because it defines the task and cannot be changed arbitrarily by the agent.
The classification follows the control boundary rather than the robot's physical outline or the agent's knowledge.
Summary
- The agent-environment boundary is based on arbitrary control, not the outer surface of a robot or animal.
- Motors, sensing hardware, muscles, and reward computation can be environmental when the agent cannot alter them arbitrarily.
- Knowledge can cross the boundary: an agent may understand the environment and its reward computation without controlling them.
- The boundary depends on the states, actions, rewards, and decision-making task being modeled.
- A policy is the action-selection rule associated with the agent, while mechanisms outside arbitrary control remain environmental.
Key Takeaways
- Use absolute control to locate the agent-environment boundary.
- Do not identify the agent by tracing a robot's shell or an animal's skin.
- Separate what the agent knows from what it can change arbitrarily.
- Keep reward computation in the environment when it defines the task and is outside the agent's arbitrary control.
- Choose the boundary in relation to the particular decision-making task.