Learning Agent and Environment
The reinforcement learning framework describes interaction between a learning agent and an environment.
The Decision Loop
When an intelligent system must act, three questions matter: What situation is it in? What can it do? What result indicates whether the outcome was desirable? Reinforcement learning organizes these questions around a learning agent and an environment. The agent learns and makes decisions through interaction rather than being described only by a final answer.
The loop is continuous in structure. The environment provides a state, the agent selects an action, and the interaction produces an environmental response, a new situation, and a reward. The same vocabulary can describe one step or many connected steps over time.
What do you think happens?
Suppose an agent is in a situation and selects an action. What should you look for next in the reinforcement learning description?
Reveal answer
Answer: The environmental response, a new situation, and a reward
The framework describes interaction through actions, environmental responses, new situations, and rewards. It keeps the decision context and the goal-related signal together.
The Framework Vocabulary
States, actions, and rewards are the three central components of the reinforcement learning framework. A state describes the situation being considered. An action describes what the agent does. A reward is the evaluative signal associated with the interaction. Rewards are numerical values arising from the environment and are maximized over time.
| Term | Question it answers | Role in the interaction |
|---|---|---|
| State | What situation is being considered? | Describes the situation available to the agent. |
| Action | What does the agent do? | Describes the agent's selected behavior. |
| Reward | What signal indicates whether the outcome was desirable? | Provides a numerical evaluative signal arising from the environment. |
The three central components connect a situation, a decision, and an evaluative signal.
The agent and environment are the participants in the framework. State, action, and reward are the vocabulary used to describe what passes between them and what the interaction means. Keeping these roles separate prevents the entire problem from becoming one vague description.
Inside and Outside the Agent
An agent is the learner and decision-maker. It learns and makes decisions through interaction. An environment is everything outside the agent that responds to its actions.
A Concrete Interaction
Consider this generated illustration: a delivery robot is the agent, and the surrounding delivery setting is the environment. At one moment, the robot's state could describe its current situation, such as being at a junction with a delivery still to complete. The robot chooses an action, such as taking one available route. The environment responds by producing a new situation and a reward that evaluates the result.
Describing One Robot Decision
Use the reinforcement learning vocabulary to describe one decision by a delivery robot.
State: Describe the robot's current situation before the decision, such as its position in the delivery setting and the fact that its task remains unfinished.
Action: Name what the robot does next, such as selecting one available route.
Environmental response: Describe the new situation produced by the interaction, such as the robot reaching another point in the setting.
Reward: Identify the numerical evaluative signal arising from the environment. The reward indicates how the outcome is treated by the task's goal.
Continuing interaction: Treat the new situation as the context for a later decision rather than stopping the description after the first action.
The situation is the state, the robot's selected route is the action, the resulting situation is the new state, and the numerical environmental signal is the reward.
Cause, Uncertainty, and Goals
The framework is useful because it preserves several features that make artificial intelligence problems difficult. It represents cause and effect by relating an action to what happens through the interaction. It allows uncertainty and nondeterminism, so a decision does not have to be treated as a perfectly predictable process. It also represents explicit goals through rewards.
Specifying a Complete Task
A reinforcement learning task is completely described when the interaction has clear roles and connections: there is an agent, an environment outside the agent, situations represented as states, choices represented as actions, environmental responses that produce new situations, and rewards that express the goal-related evaluation. This description keeps the decision context, the agent's behavior, and the goal-related signal together.
When describing a task, begin with the boundary: identify the learner and decision-maker as the agent, then classify everything outside it that responds to actions as the environment. Next name the states, actions, environmental responses, and rewards. Finally explain that the rewards are the signal being maximized over time.
Describe a generated reinforcement learning situation of your own in four parts: identify the agent and environment, name one state, name one action, and describe the new situation and reward produced by the interaction.
Hints
- Use state for the situation before the decision.
- Use action for what the agent does.
- Keep the environment separate from the agent.
- Include the reward as a numerical evaluative signal and connect it to the task's goal.
Common Description Errors
Treating the agent and environment as the same thing
The framework distinguishes the learner and decision-maker from everything outside it that responds to its actions.
Fix:
Name the agent first, then classify the surrounding responding part as the environment.Listing only an action
The framework connects the agent's situation, chosen action, environmental response, new situation, and associated reward.
Fix:
Describe the action in relation to a state, its environmental response, and its reward.Treating reward as the action's immediate result
An action is the agent's selected behavior, while a reward is a numerical evaluative signal arising from the environment.
Fix:
Keep the selected action and the environment's reward separate.Assuming every action has a perfectly predictable outcome
The framework allows uncertainty and nondeterminism in artificial intelligence problems.
Fix:
Allow the environmental response and associated reward to be part of an uncertain interaction.Describing only a final answer
The framework preserves decision context, agent behavior, cause and effect, and the goal-related signal.
Fix:
Represent the continuing interaction using states, actions, environmental responses, and rewards.
Key Takeaways
- The agent is the learner and decision-maker; the environment is everything outside the agent that responds to its actions.
- States describe situations, actions describe what the agent does, and rewards provide numerical evaluative signals.
- The interaction cycles through a state, an action, an environmental response, a new situation, and a reward.
- The framework represents cause and effect, uncertainty and nondeterminism, and explicit goals through rewards.
- A complete task description connects the agent, environment, states, actions, responses, rewards, and the goal of maximizing rewards over time.
Key Takeaways
- Reinforcement learning describes interaction between a learning agent and an environment.
- States, actions, and rewards are the framework's three central components.
- Actions lead to environmental responses, new situations, and rewards within a continuing interaction cycle.
- The framework preserves cause and effect, uncertainty and nondeterminism, and explicit goals.
- A task is fully described by connecting its agent, environment, states, actions, responses, rewards, and goal.