Policies and Action Selection
A reinforcement learning problem is an interaction between an agent and an environment.
Learning Through Interaction
A reinforcement learning problem is an interaction between an agent and an environment. The agent has a goal, but it is not given a complete recipe for reaching that goal. Instead, it learns through repeated interaction: the agent uses a state to select an action, and the environment responds with a new situation and a reward signal.
The central idea is not a single decision. It is a repeated loop in which states support decisions and rewards help evaluate them.
The Repeating Interaction
At each discrete time step, the agent receives or uses a state to decide what to do. It selects an action according to its policy. The environment then participates by providing the next situation and a reward. The next state becomes the basis for a later choice, so the interaction continues as a loop rather than ending after one action.
From State to Action
A policy describes how actions are selected from states. It connects the state used for decision-making with the action chosen by the agent.
Tracing One Task
A Generated State-Action-Reward Trace
Consider a generated reinforcement learning task in which an agent repeatedly interacts with an environment. The task has two possible situations, two possible actions, and reward signals supplied by the environment.
Identify the state: At one discrete time step, the agent is given a current situation called State A. This state is the basis for the agent's decision.
Apply the policy: The policy connects State A to Action A. The agent selects Action A rather than choosing without reference to the state.
Receive the environment response: The environment provides a new situation called State B and a reward signal. The reward helps evaluate the preceding choice.
Continue the loop: At the next discrete time step, State B supports another action selection. The new decision is part of the continuing interaction.
The trace contains the essential interaction pattern: state, action selected through a policy, next state, and reward. Repeating this pattern creates a longer effort toward the agent's goal.
The Task Interface
A particular reinforcement learning task is determined by the interface between the agent and the environment. The available actions describe what the agent can select. The states describe the situations used for decision-making. The rewards provide the signals used to evaluate those decisions. Changing this interface changes the task being specified.
| Element | Role in the task |
|---|---|
| State | Provides the situation used for a decision |
| Action | Represents a choice selected by the agent |
| Reward | Helps evaluate the choice |
The actions, states, and rewards form the task interface.
Goal and Boundary
The agent's objective is to receive as much reward as possible over time. That objective is different from the information and control available at any one point in the interaction. The agent uses a state to choose an action, while the environment provides the next situation and the reward signal. The agent therefore has a goal, but it does not receive a complete recipe for reaching that goal.
Common Reasoning Errors
Treating reinforcement learning as a one-time action
The interaction repeats across discrete time steps. A new situation and reward follow the action.
Fix:
Trace the loop: state, selected action, next situation, reward, and then the next decision.Calling the policy the reward
A policy describes how actions are selected from states. The reward helps evaluate the selected choice.
Fix:
Keep selection and evaluation separate: policy selects, reward evaluates.Assuming the agent receives a complete solution
Reinforcement learning begins with a goal but does not give the learner a complete recipe for reaching it.
Fix:
Describe learning as interaction with something outside the agent.Leaving out the task interface
The available actions, states, and rewards determine the particular task being specified.
Fix:
List the actions, states, and rewards whenever you identify a reinforcement learning task.
Check Your Understanding
A system receives a state, selects an action using a policy, and then receives a new situation and a reward. Identify the agent, the environment, the action, the next state, and the reward in this interaction. Then explain which part describes the agent's objective and which parts describe the available interaction.
Hints
- The agent is the participant that selects the action.
- The environment provides the new situation and reward.
- The objective concerns receiving as much reward as possible over time.
What do you think happens?
If the current state changes, should the action-selection step still refer to the state?
Reveal answer
Answer: Yes, because the state supports the agent's choice.
A policy describes how actions are selected from states, so the state remains part of the decision process at each discrete time step.
Key Takeaways
- A reinforcement learning problem is an interaction between an agent and an environment.
- At each discrete time step, a state supports action selection, and the environment provides a next situation and a reward.
- A policy describes how actions are selected from states.
- The available actions, states, and rewards define the task interface.
- The agent aims to receive as much reward as possible, but it learns through interaction rather than receiving a complete recipe.
Key Takeaways
- Reinforcement learning is a repeated interaction between an agent and an environment.
- States support decisions, actions are selected by the agent, and rewards help evaluate those decisions.
- A policy connects states to selected actions.
- Actions, states, and rewards together specify the task interface.
- The agent pursues reward but does not directly control the environment's responses.