Concepts / The Agent-Environment Interface for Rewards

The Agent-Environment Interface for Rewards

The learning-agent boundary follows control, not the outer boundary of a physical body.

  • Programming

A Reward Is Not the Agent's Declaration

A first impression can be tempting: if the reward expresses what the learning agent is trying to achieve, perhaps the agent should calculate its own reward. In reinforcement learning, the official reward is instead computed outside the learning agent, by the environment. The reason is that the agent should not be able to decree that it has received the reward.

The environment is the source of the official reward because the reward must remain external to the agent's own control.

Finding the Agent Boundary

The learning-agent boundary follows control, not the outer boundary of a physical body. To locate the boundary, ask: what can the learning agent control? Components within the physical body are not automatically part of the learning agent. A physically internal component can belong on the environment side of the model when it lies beyond the agent's control.

controlscontainsLearning agentcontrols limb positionsLimb positionscontrollableEnvironmentincludes what lies beyondagent controlEnergy reservesnot directly controllable
Which parts of a robot belong to the learning agent when the agent can control limb positions but cannot directly control its energy reserves?

Classifying Robot Components

A robot has energy reserves and limb positions. The learning agent can control the limb positions but cannot directly control its energy reserves. Which side of the learning-agent model should contain each component?

Check limb positions: The learning agent can control limb positions, so these positions are placed on the agent side of the model.

Check energy reserves: The energy reserves are physically inside the robot, but the learning agent cannot directly control them. They therefore remain on the environment side of the model.

Ignore the outer body boundary: The physical fact that both components are inside the robot does not decide the learning-agent boundary. Control decides it.

Limb positions are agent-side because they are controllable. Energy reserves are environment-side because they are beyond the learning agent's direct control.

Tracing the Reward Source

The reward's position follows the same boundary rule. The environment computes the official reward from the interaction that lies outside the learning agent's control. The reward then serves as information supplied to the agent, rather than a result the agent can simply announce for itself.

is evaluated bycomputesis supplied toRobot interactionoutside agent controlEnvironmentcomputes official rewardOfficial rewardexternal signalLearning agentcannot decree reward
How does information about the robot's interaction with the world flow from the environment to the learning agent as an official reward?

The official reward is environment-side even when it describes what the agent is trying to achieve.

External and Internal Rewards

Reward kindWho defines or computes itPosition in the model
Official environment rewardThe environment computes itExternal to the learning agent
Internal rewardThe agent may define itInside the learning agent
Official rewardcomputed by environmentInternal rewarddefined by agent
What is the difference between a reward calculated by the environment and a reward rule defined inside the learning agent?

Applying the Control Test

Use the control test whenever physical location creates confusion. In the robot example, energy reserves and limb positions may both be physically internal. That shared physical location does not place both inside the learning agent. Limb positions are agent-side because the agent controls them. Energy reserves are environment-side because they lie beyond the agent's direct control.

containscontainsclassified by controlclassified by controlPhysical robotcontains both componentsLimb positionscontrolled by agentAgent sidelimb positionsEnergy reservesbeyond agent controlEnvironment sideenergy reserves
Why are limb positions agent-side while energy reserves remain environment-side even though both may be physically inside the robot?

When drawing or interpreting an agent-environment diagram, label the separation as a separation of roles. Do not assume that the diagram separates physical objects. A component can remain physically inside the robot while appearing on the environment side of the learning-agent model.

Common Boundary Mistakes

  • Putting the reward inside the agent because it expresses the agent's goal.

    The official reward is computed outside the agent. The agent should not be able to decree that it has received the reward.

    Fix: Place the official reward source in the environment, even though the reward describes what the agent is trying to achieve.

  • Using the robot's outer physical boundary as the learning-agent boundary.

    The learning-agent boundary follows control rather than physical location.

    Fix: Ask which components the learning agent can control.

  • Treating an internal reward as identical to the environment's reward.

    Internal rewards are allowed, but they are distinct from the environment's reward.

    Fix: Record who defines or computes the reward before classifying it.

  • Classifying both energy reserves and limb positions in the same way merely because both are inside the robot.

    Limb positions can be agent-side when the agent controls them, while energy reserves can remain environment-side when they are beyond the agent's control.

    Fix: Apply the control test separately to each component.

Boundary Practice

EASY

A physical robot contains a component whose behavior is beyond the learning agent's control and another component whose position the learning agent can control. Classify each component as agent-side or environment-side. Then state which system should compute the official reward.

Hints
  • Do not begin with physical location.
  • Ask what the learning agent can control.
  • The official reward must come from outside the learning agent.

Practice Check

Classify a controllable limb position and an uncontrollable energy reserve, then identify the official reward source.

Limb position: Because the learning agent controls the limb position, classify it on the agent side.

Energy reserve: Because the learning agent cannot directly control the energy reserve, classify it on the environment side.

Official reward: The environment computes the official reward, so the agent cannot simply declare that it has received it.

Control determines the boundary: limb position is agent-side, energy reserve is environment-side, and the environment computes the official reward.

Key Takeaways

  1. The learning-agent boundary follows control, not the physical outer boundary of the robot.
  2. The environment computes the official reward because the agent should not be able to decree that it has received it.
  3. A physically internal component can still be environment-side when it lies beyond the learning agent's control.
  4. Internal rewards may be defined by the agent, but they remain distinct from the environment's reward.
  5. For the robot example, controllable limb positions are agent-side while uncontrollable energy reserves are environment-side.

Key Takeaways

  • The reward source belongs outside the learning agent because the agent must not be able to declare its own official reward.
  • The correct boundary question is what the learning agent can control.
  • Physical containment and learning-agent membership are not automatically the same.
  • Internal rewards are possible, but they must be distinguished from the environment-computed reward.
  • Robot limb positions and energy reserves are classified differently when only the limb positions are under the agent's control.