Reinforcement Learning Agents and Environments
Do not identify the agent by tracing the outside of a robot's or animal's body.
Beyond the Physical Shell
When people first imagine a reinforcement learning agent, they often draw a boundary around everything inside a robot's shell or an animal's skin. That is usually too broad. In reinforcement learning, the important boundary is functional: it depends on what the agent can control, not on where the physical body ends.
A physical component can belong to the robot or animal physically while still being treated as part of the environment. The agent-environment boundary can therefore cut across the physical body.
The Absolute-Control Test
Use absolute control as the main test for classifying a component. Ask whether the agent can control that component directly and change its behavior at will. Components under the agent's absolute control belong inside the agent's functional boundary. Components that the agent cannot alter arbitrarily remain part of the environment, even if they are physically located within the same robot or animal.
Applying the Absolute-Control Test
An imaginary learning system has an action-selection component, a component that computes rewards, and a physical component inside the same machine. Classify each component using absolute control.
Action selection: Place this on the agent side because it represents the agent's choice of action.
Reward computation: Place this on the environment side because the reward calculation defines the task and cannot be changed arbitrarily by the agent.
Physical component: Do not classify it from its physical location alone. Ask whether the agent has absolute control over it. If it cannot alter the component at will, treat it as part of the environment.
The classification follows control, not the machine's outer shell.
Reward Comes from Outside
Reward computation belongs to the environment because it defines the task. The agent may understand exactly how rewards are computed from actions and states, but understanding the rule does not give the agent the power to rewrite that rule arbitrarily. The reward signal is therefore treated as coming from outside the agent.
Knowledge can cross the agent-environment boundary. Absolute control defines the boundary, so an agent can know the reward rule without owning or changing the reward computation.
One Interaction Cycle
The interaction can be understood as a repeating exchange. The agent selects an action. The environment responds with a new observation and a reward signal. The reward is still external because the environment's computation defines the task, even if the agent understands the computation completely.
What do you think happens?
Suppose the agent knows the exact rule used to calculate rewards. Does that knowledge move reward computation inside the agent?
Reveal answer
Answer: No, because knowledge does not imply absolute control.
The reward computation remains external when the agent cannot change the task rule arbitrarily. The boundary is determined by control, not by the limit of the agent's knowledge.
Control and Knowledge
The limit of control and the limit of knowledge are different. An agent may understand its environment completely, including how rewards are calculated from actions and states, while still lacking the ability to alter those rules. Consequently, a difficult task does not require the environment to be mysterious or surprising. The difficulty can remain in finding actions that accomplish the task under known rules.
Common Boundary Mistakes
Treating everything inside a robot's shell or an animal's skin as the agent.
The physical outline is not the decisive boundary. A component inside the body can still be part of the environment.
Fix:
Apply the absolute-control test instead of tracing the body's outside.Treating reward computation as part of the agent because the agent knows the reward rule.
Understanding a rule does not mean the agent can change it arbitrarily.
Fix:
Keep reward computation external when it defines the task and lies outside the agent's absolute control.Assuming a task must be difficult because the agent is ignorant of the environment.
Difficulty can remain in finding actions that accomplish the task even when the rules are understood.
Fix:
Separate the challenge of choosing effective actions from the question of how much the agent knows.
Boundary Classification Practice
For each item in this imaginary learning system, decide whether it belongs inside the agent's functional boundary or remains part of the environment: the action-selection process, the reward computation, and a physical component that the agent cannot change arbitrarily. Explain each decision using absolute control rather than physical location or knowledge.
Hints
- Start by asking which item represents the agent's choice.
- Ask whether the agent can alter the reward rule at will.
- Do not use the item's location inside or outside a machine as your main test.
Checking Your Classification
Classify the three items using the rule of absolute control.
Action-selection process: Classify it on the agent side because it represents the agent selecting an action.
Reward computation: Classify it on the environment side because it defines the task and cannot be changed arbitrarily by the agent.
Unchangeable physical component: Classify it on the environment side because its physical location does not override the fact that the agent lacks absolute control.
The correct classifications follow control: action selection is associated with the agent, while reward computation and the specified uncontrolled physical component remain external.
Key Takeaways
- The agent-environment boundary is functional rather than a copy of a robot's shell or an animal's skin.
- Absolute control is the main test for deciding whether something belongs to the agent.
- Reward computation remains external because it defines the task and cannot be changed arbitrarily by the agent.
- An agent can understand the environment, including reward rules, without controlling those rules.
- A task can be difficult because effective actions are hard to find, even when the environment is completely understood.
Key Takeaways
- The agent is not necessarily everything inside a robot or animal.
- Use absolute control, not physical location, to classify components.
- Reward computation belongs to the environment because it defines the task.
- Knowledge may cross the agent-environment boundary, but absolute control determines that boundary.
- Complete understanding does not remove the difficulty of selecting actions that accomplish the task.