Concepts / Learning Agent and Environment

Learning Agent and Environment

The reinforcement learning framework describes interaction between a learning agent and an environment.

  • Programming

The Decision Loop

When an intelligent system must act, three questions matter: What situation is it in? What can it do? What result indicates whether the outcome was desirable? Reinforcement learning organizes these questions around a learning agent and an environment. The agent learns and makes decisions through interaction rather than being described only by a final answer.

stateactionnew situation and rewardAgentlearns and decidesEnvironmentresponds to actions
What happens as an agent observes a situation, acts, receives an environmental response and reward, and encounters a new situation?

The loop is continuous in structure. The environment provides a state, the agent selects an action, and the interaction produces an environmental response, a new situation, and a reward. The same vocabulary can describe one step or many connected steps over time.

What do you think happens?

Suppose an agent is in a situation and selects an action. What should you look for next in the reinforcement learning description?

  • Only the agent's internal thoughts
  • The environmental response, a new situation, and a reward
  • Only the final answer to the entire task
Reveal answer

Answer: The environmental response, a new situation, and a reward

The framework describes interaction through actions, environmental responses, new situations, and rewards. It keeps the decision context and the goal-related signal together.

The Framework Vocabulary

States, actions, and rewards are the three central components of the reinforcement learning framework. A state describes the situation being considered. An action describes what the agent does. A reward is the evaluative signal associated with the interaction. Rewards are numerical values arising from the environment and are maximized over time.

considered by agentenvironment respondsinteraction producesCurrent statesituationActionagent choiceNew statenew situationRewardnumerical signal
How do a current state, a selected action, the resulting new state, and the reward connect in one step?
TermQuestion it answersRole in the interaction
StateWhat situation is being considered?Describes the situation available to the agent.
ActionWhat does the agent do?Describes the agent's selected behavior.
RewardWhat signal indicates whether the outcome was desirable?Provides a numerical evaluative signal arising from the environment.

The three central components connect a situation, a decision, and an evaluative signal.

The agent and environment are the participants in the framework. State, action, and reward are the vocabulary used to describe what passes between them and what the interaction means. Keeping these roles separate prevents the entire problem from becoming one vague description.

Inside and Outside the Agent

An agent is the learner and decision-maker. It learns and makes decisions through interaction. An environment is everything outside the agent that responds to its actions.

providesdescribes situationselectsaffects interactionproducesevaluates outcomeAgentlearner and decision-makerStatesituationEnvironmenteverything outside theagentActionagent choiceRewardnumerical signal
Which participants are inside the agent, which are outside it, and how do states, actions, and rewards cross the boundary?

A Concrete Interaction

Consider this generated illustration: a delivery robot is the agent, and the surrounding delivery setting is the environment. At one moment, the robot's state could describe its current situation, such as being at a junction with a delivery still to complete. The robot chooses an action, such as taking one available route. The environment responds by producing a new situation and a reward that evaluates the result.

Describing One Robot Decision

Use the reinforcement learning vocabulary to describe one decision by a delivery robot.

State: Describe the robot's current situation before the decision, such as its position in the delivery setting and the fact that its task remains unfinished.

Action: Name what the robot does next, such as selecting one available route.

Environmental response: Describe the new situation produced by the interaction, such as the robot reaching another point in the setting.

Reward: Identify the numerical evaluative signal arising from the environment. The reward indicates how the outcome is treated by the task's goal.

Continuing interaction: Treat the new situation as the context for a later decision rather than stopping the description after the first action.

The situation is the state, the robot's selected route is the action, the resulting situation is the new state, and the numerical environmental signal is the reward.

agent selectsenvironment respondsinteraction producesJunctioncurrent situationRoute choiceselected actionNext pointnew situationRewardenvironmental signal
How can a simple generated situation be described using state, action, new state, and reward?

Cause, Uncertainty, and Goals

The framework is useful because it preserves several features that make artificial intelligence problems difficult. It represents cause and effect by relating an action to what happens through the interaction. It allows uncertainty and nondeterminism, so a decision does not have to be treated as a perfectly predictable process. It also represents explicit goals through rewards.

agent choosesmay producemay produceassociated withassociated withStatecurrent situationReward Aevaluative signalActionselected decisionReward Bevaluative signalOutcome Aenvironmental responseOutcome Benvironmental response
How can the same kind of action in the same kind of situation be connected to different possible environmental outcomes or rewards?
agent considerscauses interactionreceives evaluationcontributes toSituationdecision contextActionagent behaviorOutcomeenvironmental resultRewardnumerical evaluationGoalmaximize over time
How do rewards connect an immediate outcome with the agent's longer-term goal?

Specifying a Complete Task

A reinforcement learning task is completely described when the interaction has clear roles and connections: there is an agent, an environment outside the agent, situations represented as states, choices represented as actions, environmental responses that produce new situations, and rewards that express the goal-related evaluation. This description keeps the decision context, the agent's behavior, and the goal-related signal together.

provides situationsselects choiceselicits interactionproduces new situationsproduces signalssupportsAgentlearner and decision-makerStatessituationsGoalmaximize rewards over timeEnvironmentoutside the agentActionsagent choicesResponsesenvironmental outcomesRewardsnumerical evaluations
What elements must be connected to specify the states, actions, environmental responses, rewards, and goal of a reinforcement learning task?

When describing a task, begin with the boundary: identify the learner and decision-maker as the agent, then classify everything outside it that responds to actions as the environment. Next name the states, actions, environmental responses, and rewards. Finally explain that the rewards are the signal being maximized over time.

EASY

Describe a generated reinforcement learning situation of your own in four parts: identify the agent and environment, name one state, name one action, and describe the new situation and reward produced by the interaction.

Hints
  • Use state for the situation before the decision.
  • Use action for what the agent does.
  • Keep the environment separate from the agent.
  • Include the reward as a numerical evaluative signal and connect it to the task's goal.

Common Description Errors

  • Treating the agent and environment as the same thing

    The framework distinguishes the learner and decision-maker from everything outside it that responds to its actions.

    Fix: Name the agent first, then classify the surrounding responding part as the environment.

  • Listing only an action

    The framework connects the agent's situation, chosen action, environmental response, new situation, and associated reward.

    Fix: Describe the action in relation to a state, its environmental response, and its reward.

  • Treating reward as the action's immediate result

    An action is the agent's selected behavior, while a reward is a numerical evaluative signal arising from the environment.

    Fix: Keep the selected action and the environment's reward separate.

  • Assuming every action has a perfectly predictable outcome

    The framework allows uncertainty and nondeterminism in artificial intelligence problems.

    Fix: Allow the environmental response and associated reward to be part of an uncertain interaction.

  • Describing only a final answer

    The framework preserves decision context, agent behavior, cause and effect, and the goal-related signal.

    Fix: Represent the continuing interaction using states, actions, environmental responses, and rewards.

Key Takeaways

  1. The agent is the learner and decision-maker; the environment is everything outside the agent that responds to its actions.
  2. States describe situations, actions describe what the agent does, and rewards provide numerical evaluative signals.
  3. The interaction cycles through a state, an action, an environmental response, a new situation, and a reward.
  4. The framework represents cause and effect, uncertainty and nondeterminism, and explicit goals through rewards.
  5. A complete task description connects the agent, environment, states, actions, responses, rewards, and the goal of maximizing rewards over time.

Key Takeaways

  • Reinforcement learning describes interaction between a learning agent and an environment.
  • States, actions, and rewards are the framework's three central components.
  • Actions lead to environmental responses, new situations, and rewards within a continuing interaction cycle.
  • The framework preserves cause and effect, uncertainty and nondeterminism, and explicit goals.
  • A task is fully described by connecting its agent, environment, states, actions, responses, rewards, and goal.