Concepts / Actions and Environmental Responses

Actions and Environmental Responses

An agent learns and makes decisions through interaction.

  • Programming

Learning Through Interaction

Reinforcement learning is about learning while interacting with something outside the learner. The learner and decision-maker is called the agent. The surrounding part that responds to the agent is called the environment. The central question is therefore not only what the agent knows, but what it should do next.

An agent learns and makes decisions through interaction, while trying to achieve a goal.

The Interaction Cycle

The interaction follows a repeating cycle. The agent is in a situation, chooses an action, and the environment responds. That response leads to a new situation and produces a reward. Rewards are numerical values arising from the environment. Over time, the agent aims to maximize the rewards it receives.

situationactionresponse, new situation, rewardAgentLearner and decision-makerEnvironmentEverything outside theagent
What happens as the agent observes a situation, chooses an action, receives an environmental response and reward, and reaches a new situation?

The diagram shows the direction of the interaction. The agent supplies an action. The environment supplies the response, the new situation, and the numerical reward. The cycle can then continue from that new situation.

Tracing One Decision

A delivery robot chooses where to move

Trace one interaction cycle for a hypothetical delivery robot.

Situation: The robot is in a particular situation, such as being at a location while carrying a delivery.

Action: The robot chooses an action, such as moving toward a delivery destination.

Environmental response: The environment responds to that action. The response may place the robot in a new situation.

Reward: The environment produces a numerical reward for the interaction.

Next cycle: The robot continues making decisions from the new situation while trying to maximize rewards over time.

One decision is not the whole learning problem. Reinforcement learning connects repeated actions, environmental responses, new situations, and rewards into an interaction cycle.

The Agent–Environment Boundary

The agent is the learner and decision-maker. The environment is everything outside the agent that responds to its actions.

choosesaffectsproducesproducesAgentLearner and decision-makerActionChosen by agentEnvironmentEverything outside theagentResponseProduced by environmentRewardNumerical value fromenvironment
Which parts of the interaction belong to the agent, and which belong to everything outside the agent?

The boundary is defined by roles in the interaction. The agent is the part that learns and makes decisions. Everything outside that agent belongs to the environment when it responds to the agent's actions. This distinction helps identify who chooses an action and where the response and reward come from.

Actions and Responses

Part of interactionSourceRole
ActionAgentA choice made by the learner and decision-maker
Environmental responseEnvironmentWhat everything outside the agent produces in response
RewardEnvironmentA numerical value arising from the environment
New situationEnvironment's responseThe situation from which the next interaction can continue
followed byproducesleads toActionChosen by agentRewardNumerical valueEnvironmentalresponseProduced by environmentNew situationNext point in the cycle
How is an action chosen by the agent different from the environmental response that follows it?

Do not treat an action and a response as the same event. The action is the agent's choice. The response is supplied by the environment after that choice.

Specifying the Task

A reinforcement learning task is fully specified when the interaction is defined clearly enough to identify its situations, the actions available to the agent, the environmental responses, the rewards, and the conditions that determine when the interaction stops. These pieces describe what the agent can encounter, what it can choose, what the environment can do in response, how outcomes are evaluated numerically, and whether another cycle remains.

specifiesspecifiesspecifiesspecifiesspecifiesReinforcementlearning taskComplete interactionspecificationSituationsWhat the agent canencounterAvailable actionsWhat the agent can chooseEnvironmentalresponsesWhat follows an actionRewardsNumerical valuesStopping conditionsWhen interaction ends
What components must be specified to fully define the situations, available actions, environmental responses, rewards, and stopping conditions of a reinforcement learning problem?

When describing a reinforcement learning problem, name both sides of the interaction. State what the agent can do and what the environment can return. Then identify the numerical reward and the point at which the interaction ends.

Mistakes in Identifying Roles

  • Treating the environment as only a physical place

    The environment is defined as everything outside the agent that responds to its actions.

    Fix: Identify the environment by its role in the interaction, not by whether it is a physical location.

  • Calling the reward the agent's action

    The reward arises from the environment, while the action is chosen by the agent.

    Fix: Keep the sequence clear: the agent chooses an action, and the environment supplies a response and reward.

  • Describing one action as the entire learning process

    The interaction cycle also includes the environmental response, a new situation, and a reward.

    Fix: Trace the cycle through the environment's response and into the next situation.

  • Leaving the task underspecified

    Those components are needed to define the reinforcement learning interaction completely.

    Fix: List each component before claiming that the task has been fully specified.

Check the Cycle

EASY

A learner says: The environment chooses an action, and the agent gives itself a reward. Rewrite this description so that it correctly identifies the agent's role, the environment's role, and the order of the interaction.

Hints
  • Start with the part that chooses the action.
  • Then identify what comes from the environment.
  • Include the new situation and numerical reward in the cycle.

A corrected interaction description

Correct the statement that the environment chooses an action and the agent gives itself a reward.

Identify the chooser: The agent is the learner and decision-maker, so it chooses the action.

Identify the responder: The environment is everything outside the agent that responds to the action.

Identify the result: The environmental response leads to a new situation and a numerical reward arising from the environment.

The agent chooses an action. The environment responds, producing a new situation and a numerical reward.

Key Takeaways

  1. The agent is the learner and decision-maker in reinforcement learning.
  2. The environment is everything outside the agent that responds to its actions.
  3. The interaction cycles through an action, an environmental response, a new situation, and a numerical reward.
  4. Rewards arise from the environment and are maximized over time.
  5. A complete task specifies its situations, available actions, environmental responses, rewards, and stopping conditions.

Key Takeaways

  • An agent learns and makes decisions through interaction.
  • The agent chooses actions, while the environment responds and produces rewards.
  • Each response leads to a new situation, allowing the interaction cycle to continue.
  • A complete reinforcement learning task specifies situations, actions, responses, rewards, and stopping conditions.
  • The environment is everything outside the agent that responds to its actions.