Concepts / Policies and Action Selection

Policies and Action Selection

A reinforcement learning problem is an interaction between an agent and an environment.

  • Programming

Learning Through Interaction

A reinforcement learning problem is an interaction between an agent and an environment. The agent has a goal, but it is not given a complete recipe for reaching that goal. Instead, it learns through repeated interaction: the agent uses a state to select an action, and the environment responds with a new situation and a reward signal.

The central idea is not a single decision. It is a repeated loop in which states support decisions and rewards help evaluate them.

The Repeating Interaction

At each discrete time step, the agent receives or uses a state to decide what to do. It selects an action according to its policy. The environment then participates by providing the next situation and a reward. The next state becomes the basis for a later choice, so the interaction continues as a loop rather than ending after one action.

supports selectionagent actsprovidesprovidessupports the next choiceStatesituation at time tActionselected at time tEnvironmentresponds to actionNext statenew situationRewardevaluation signal
What happens next as the agent observes a state, selects an action, and receives a new state and reward from the environment?

From State to Action

A policy describes how actions are selected from states. It connects the state used for decision-making with the action chosen by the agent.

provided toprovided toselectsselectsState Adecision inputPolicyselects actionsAction Aselected responseState Bdecision inputAction Bselected response
How does a policy connect each observed state to the action the agent selects?

Tracing One Task

A Generated State-Action-Reward Trace

Consider a generated reinforcement learning task in which an agent repeatedly interacts with an environment. The task has two possible situations, two possible actions, and reward signals supplied by the environment.

Identify the state: At one discrete time step, the agent is given a current situation called State A. This state is the basis for the agent's decision.

Apply the policy: The policy connects State A to Action A. The agent selects Action A rather than choosing without reference to the state.

Receive the environment response: The environment provides a new situation called State B and a reward signal. The reward helps evaluate the preceding choice.

Continue the loop: At the next discrete time step, State B supports another action selection. The new decision is part of the continuing interaction.

The trace contains the essential interaction pattern: state, action selected through a policy, next state, and reward. Repeating this pattern creates a longer effort toward the agent's goal.

providesselectssent toprovidessupports choicehelps evaluate choiceAgentselects actionsStatebasis for choiceEnvironmentprovides responsesActionselected responseRewardchoice evaluation
What does each element represent, and how are these elements connected during a reinforcement learning interaction?

The Task Interface

A particular reinforcement learning task is determined by the interface between the agent and the environment. The available actions describe what the agent can select. The states describe the situations used for decision-making. The rewards provide the signals used to evaluate those decisions. Changing this interface changes the task being specified.

helps specifyhelps specifyhelps specifyselects fromprovidesprovidesReinforcementlearning taskspecified by an interfaceStatesdecision situationsAgentchooses actionsActionsavailable choicesEnvironmentprovides situations andrewardsRewardsevaluation signals
How do the possible states, available actions, and reward signals work together to specify a particular reinforcement learning task?
ElementRole in the task
StateProvides the situation used for a decision
ActionRepresents a choice selected by the agent
RewardHelps evaluate the choice

The actions, states, and rewards form the task interface.

Goal and Boundary

The agent's objective is to receive as much reward as possible over time. That objective is different from the information and control available at any one point in the interaction. The agent uses a state to choose an action, while the environment provides the next situation and the reward signal. The agent therefore has a goal, but it does not receive a complete recipe for reaching that goal.

usesselectsagent acts throughprovidesprovidessupports selectionAgentselects actionsPolicyconnects states to actionsActionagent choiceEnvironmentoutside the agentStatedecision informationRewardevaluation signal
Which information and controls are inside the agent, and which states, rewards, and dynamics belong to the environment?
pursues throughsupportsleads to environment responsearrives withReceive more rewardagent objectiveStatebasis for choiceActionagent controlRewardevaluation signalNext stateenvironment response
How is the agent's goal different from what the agent can observe and the actions it can control?

Common Reasoning Errors

  • Treating reinforcement learning as a one-time action

    The interaction repeats across discrete time steps. A new situation and reward follow the action.

    Fix: Trace the loop: state, selected action, next situation, reward, and then the next decision.

  • Calling the policy the reward

    A policy describes how actions are selected from states. The reward helps evaluate the selected choice.

    Fix: Keep selection and evaluation separate: policy selects, reward evaluates.

  • Assuming the agent receives a complete solution

    Reinforcement learning begins with a goal but does not give the learner a complete recipe for reaching it.

    Fix: Describe learning as interaction with something outside the agent.

  • Leaving out the task interface

    The available actions, states, and rewards determine the particular task being specified.

    Fix: List the actions, states, and rewards whenever you identify a reinforcement learning task.

Check Your Understanding

EASY

A system receives a state, selects an action using a policy, and then receives a new situation and a reward. Identify the agent, the environment, the action, the next state, and the reward in this interaction. Then explain which part describes the agent's objective and which parts describe the available interaction.

Hints
  • The agent is the participant that selects the action.
  • The environment provides the new situation and reward.
  • The objective concerns receiving as much reward as possible over time.

What do you think happens?

If the current state changes, should the action-selection step still refer to the state?

  • Yes, because the state supports the agent's choice
  • No, because actions are selected independently of states
Reveal answer

Answer: Yes, because the state supports the agent's choice.

A policy describes how actions are selected from states, so the state remains part of the decision process at each discrete time step.

Key Takeaways

  1. A reinforcement learning problem is an interaction between an agent and an environment.
  2. At each discrete time step, a state supports action selection, and the environment provides a next situation and a reward.
  3. A policy describes how actions are selected from states.
  4. The available actions, states, and rewards define the task interface.
  5. The agent aims to receive as much reward as possible, but it learns through interaction rather than receiving a complete recipe.

Key Takeaways

  • Reinforcement learning is a repeated interaction between an agent and an environment.
  • States support decisions, actions are selected by the agent, and rewards help evaluate those decisions.
  • A policy connects states to selected actions.
  • Actions, states, and rewards together specify the task interface.
  • The agent pursues reward but does not directly control the environment's responses.