Concepts / Defining Goals and Rewards in Reinforcement Learning

Defining Goals and Rewards in Reinforcement Learning

The learning-agent boundary follows control, not the outer boundary of a physical body.

  • Programming

The Boundary Question

A reward expresses what a reinforcement-learning system is trying to achieve, so it may seem natural for the learning agent to calculate its own reward. The important modeling rule is different: the official reward is computed outside the agent, by the environment. To decide where the agent ends and the environment begins, ask what the learning agent can control. The boundary follows control, not the outer surface of a physical body.

controlslies withinreturns rewardLearning agentcontrol sourceLimb positionsagent-controlledEnergy reservesbeyond agent controlEnvironmentcomputes official reward
Which parts of the robot can the agent control, and where should the agent–environment boundary be drawn?

Physical location is not enough to identify the agent. A component can be inside a robot's body and still be on the environment side of the learning-agent model if it lies beyond the agent's control.

Tracing the Reward Source

A Robot Interaction

Determine where the official reward belongs when a robot's learning agent controls limb positions but does not control its energy reserves.

1. Locate control: The learning agent controls the robot's limb positions. Those positions therefore lie on the agent side of the model.

2. Check the energy reserves: The energy reserves are physically inside the robot, but they are beyond the learning agent's control in this model. They therefore lie on the environment side.

3. Place the reward source: The environment uses the interaction outside the agent to compute the official reward. The agent receives that reward; it does not decree that it has received it.

The control boundary places limb positions with the agent, energy reserves with the environment, and the official reward computation with the environment.

producesinformsproducesAgent behaviorcontrolled limb positionsEnvironmentobservationinteraction resultReward computationenvironment-sideOfficial rewardreturned to agent
How does the environment observe the agent's behavior, compute a reward, and send that reward back to the agent?

The separation protects the meaning of the official reward. If the agent could simply decree that it had received the reward, the reward would no longer be an external signal describing the result of the agent's interaction with the environment. The environment is therefore assigned the role of computing and returning that official signal.

Physical Inside Versus Agent Inside

The word internal can create confusion. It can mean physically inside the robot, or it can mean inside the learning agent. Those meanings are not automatically the same. Reward modeling uses the second meaning: what lies within the learning agent's control. A robot's energy reserves may be located inside its body while remaining environment-side because the learning agent does not control them. By contrast, limb positions belong on the agent side when the agent controls them.

agent controlnot agent controlLimb positionsinitial positionsLimb positionsnew positionsEnergy reservesenvironment-sideEnergy reservesenvironment-side
How do agent actions change limb positions, and how is that different from an environment variable such as energy reserves?

External and Internal Signals

Signal typeWhere it is definedRelationship to the official reward
Environment rewardOutside the learning agent, in the environmentIt is the official reward returned to the agent
Internal rewardInside the learning agentIt is allowed, but remains distinct from the environment's reward

An agent may define an internal reward, but that does not move the official reward computation inside the agent. The two ideas must remain separate: an externally computed reward belongs to the environment, while an internally defined reward is a signal created by the agent. Calling a signal internal does not by itself tell you whether it is physically inside the robot; it tells you that it is within the learning agent.

definesdefinesEnvironment rewardofficial signalEnvironmentcomputes rewardInternal rewardagent-defined signalLearning agentdefines internal signal
What is the difference between a reward computed by the environment and a goal or signal defined within the agent?

Boundary Practice

EASY

A robot contains an energy reservoir and movable limbs. In a proposed learning-agent model, the agent controls the limb positions but does not control the energy reservoir. Classify each component as agent-side or environment-side, then identify who computes the official reward.

Hints
  • Ask what the learning agent can control rather than where each component is physically located.
  • The official reward is not defined by the agent merely because it expresses what the agent is trying to achieve.

Practice Check

Use the control rule to classify the robot's limbs, energy reserves, and official reward.

Limb positions: They are agent-side because the learning agent controls them.

Energy reserves: They are environment-side because they lie beyond the agent's control, even though they are physically inside the robot.

Official reward: It is environment-side because the environment computes it and sends it to the agent.

The model separates roles by control: the agent controls limb positions, the environment includes the uncontrolled energy reserves, and the environment computes the official reward.

Common Boundary Mistakes

  • Placing the reward inside the agent because the reward describes the agent's goal.

    The official reward is computed outside the agent. The agent should not be able to decree that it has received it.

    Fix: Place the official reward computation in the environment and treat any agent-defined signal as an internal reward distinct from it.

  • Using the robot's physical boundary as the learning-agent boundary.

    The model separates roles by control, not by the outer boundary of a physical body.

    Fix: Ask whether the learning agent controls the component. Energy reserves can be environment-side even when physically inside the robot.

  • Assuming that every physically internal component is agent-internal.

    Physical internality and being inside the learning agent refer to different systems.

    Fix: Use internal in the learning-agent sense only after checking the control boundary.

  • Treating an internal reward as the same thing as the environment's reward.

    Internal rewards are allowed, but they remain distinct from the environment's reward.

    Fix: Label the internally defined signal separately and keep the environment responsible for the official reward.

The Control Rule

  1. The learning-agent boundary follows control, not the physical outer boundary of a robot.
  2. The environment computes the official reward because the agent should not be able to decree that it has received it.
  3. Energy reserves can be environment-side even when they are physically inside the robot.
  4. Limb positions are agent-side when the learning agent controls them.
  5. An internal reward may be defined by the agent, but it remains distinct from the environment's official reward.

Key Takeaways

  • Draw the agent–environment boundary where the learning agent's control stops.
  • Do not use physical location alone to decide whether something belongs to the agent.
  • Assign the official reward to the environment so the agent cannot simply decree that it received the reward.
  • Keep internally defined rewards conceptually separate from the environment's reward.
  • For the robot example, limb positions are agent-side and energy reserves are environment-side when only the former are controlled by the agent.