The Agent-Environment Interface for Rewards
The learning-agent boundary follows control, not the outer boundary of a physical body.
A Reward Is Not the Agent's Declaration
A first impression can be tempting: if the reward expresses what the learning agent is trying to achieve, perhaps the agent should calculate its own reward. In reinforcement learning, the official reward is instead computed outside the learning agent, by the environment. The reason is that the agent should not be able to decree that it has received the reward.
The environment is the source of the official reward because the reward must remain external to the agent's own control.
Finding the Agent Boundary
The learning-agent boundary follows control, not the outer boundary of a physical body. To locate the boundary, ask: what can the learning agent control? Components within the physical body are not automatically part of the learning agent. A physically internal component can belong on the environment side of the model when it lies beyond the agent's control.
Classifying Robot Components
A robot has energy reserves and limb positions. The learning agent can control the limb positions but cannot directly control its energy reserves. Which side of the learning-agent model should contain each component?
Check limb positions: The learning agent can control limb positions, so these positions are placed on the agent side of the model.
Check energy reserves: The energy reserves are physically inside the robot, but the learning agent cannot directly control them. They therefore remain on the environment side of the model.
Ignore the outer body boundary: The physical fact that both components are inside the robot does not decide the learning-agent boundary. Control decides it.
Limb positions are agent-side because they are controllable. Energy reserves are environment-side because they are beyond the learning agent's direct control.
Tracing the Reward Source
The reward's position follows the same boundary rule. The environment computes the official reward from the interaction that lies outside the learning agent's control. The reward then serves as information supplied to the agent, rather than a result the agent can simply announce for itself.
The official reward is environment-side even when it describes what the agent is trying to achieve.
External and Internal Rewards
| Reward kind | Who defines or computes it | Position in the model |
|---|---|---|
| Official environment reward | The environment computes it | External to the learning agent |
| Internal reward | The agent may define it | Inside the learning agent |
Applying the Control Test
Use the control test whenever physical location creates confusion. In the robot example, energy reserves and limb positions may both be physically internal. That shared physical location does not place both inside the learning agent. Limb positions are agent-side because the agent controls them. Energy reserves are environment-side because they lie beyond the agent's direct control.
When drawing or interpreting an agent-environment diagram, label the separation as a separation of roles. Do not assume that the diagram separates physical objects. A component can remain physically inside the robot while appearing on the environment side of the learning-agent model.
Common Boundary Mistakes
Putting the reward inside the agent because it expresses the agent's goal.
The official reward is computed outside the agent. The agent should not be able to decree that it has received the reward.
Fix:
Place the official reward source in the environment, even though the reward describes what the agent is trying to achieve.Using the robot's outer physical boundary as the learning-agent boundary.
The learning-agent boundary follows control rather than physical location.
Fix:
Ask which components the learning agent can control.Treating an internal reward as identical to the environment's reward.
Internal rewards are allowed, but they are distinct from the environment's reward.
Fix:
Record who defines or computes the reward before classifying it.Classifying both energy reserves and limb positions in the same way merely because both are inside the robot.
Limb positions can be agent-side when the agent controls them, while energy reserves can remain environment-side when they are beyond the agent's control.
Fix:
Apply the control test separately to each component.
Boundary Practice
A physical robot contains a component whose behavior is beyond the learning agent's control and another component whose position the learning agent can control. Classify each component as agent-side or environment-side. Then state which system should compute the official reward.
Hints
- Do not begin with physical location.
- Ask what the learning agent can control.
- The official reward must come from outside the learning agent.
Practice Check
Classify a controllable limb position and an uncontrollable energy reserve, then identify the official reward source.
Limb position: Because the learning agent controls the limb position, classify it on the agent side.
Energy reserve: Because the learning agent cannot directly control the energy reserve, classify it on the environment side.
Official reward: The environment computes the official reward, so the agent cannot simply declare that it has received it.
Control determines the boundary: limb position is agent-side, energy reserve is environment-side, and the environment computes the official reward.
Key Takeaways
- The learning-agent boundary follows control, not the physical outer boundary of the robot.
- The environment computes the official reward because the agent should not be able to decree that it has received it.
- A physically internal component can still be environment-side when it lies beyond the learning agent's control.
- Internal rewards may be defined by the agent, but they remain distinct from the environment's reward.
- For the robot example, controllable limb positions are agent-side while uncontrollable energy reserves are environment-side.
Key Takeaways
- The reward source belongs outside the learning agent because the agent must not be able to declare its own official reward.
- The correct boundary question is what the learning agent can control.
- Physical containment and learning-agent membership are not automatically the same.
- Internal rewards are possible, but they must be distinguished from the environment-computed reward.
- Robot limb positions and energy reserves are classified differently when only the limb positions are under the agent's control.