The Reinforcement Learning Agent
A policy connects each perceived state with an action.
From Situation to Behavior
An agent cannot behave until its perception of the environment is connected to an action. A policy provides that connection. It describes what the agent does when it encounters a perceived state, so studying a policy means studying the rule that turns a situation into behavior.
A policy connects each perceived state with an action. Following those connections gives the agent's behavior.
How a Policy Selects an Action
The policy is the rule that answers this question: given the situation the agent perceives, what action should it take? The perceived situation is the input to the policy, and the selected action is its behavioral result. A policy may be represented as a simple function or lookup table, or producing the action may involve extensive computation, such as a search process. These forms differ in how the mapping is carried out, but the policy's behavioral role remains the same.
A Perceived State Becomes an Action
Suppose an agent has a policy that connects the perceived state "target ahead" with the action "move toward target." What does the policy do when the agent perceives that state?
Perceive: The agent encounters the perceived state "target ahead."
Apply the policy: The policy looks up or computes the action connected to that perceived state.
Act: The connected action is "move toward target."
The policy turns the perceived state "target ahead" into the action "move toward target."
Consistent and Randomized Choices
| Policy behavior | What happens for a perceived state | Behavioral result |
|---|---|---|
| Consistent selection | The policy selects an action consistently | The same state is connected with the same selected action |
| Stochastic selection | The policy can involve randomness when selecting among actions | The same state can lead to different action selections |
The Agent–Environment Framework
The reinforcement learning framework describes interaction between a learning agent and an environment. The agent and environment are the two participants. States, actions, and rewards are the vocabulary used to describe their interaction: the state describes the situation being considered, the action describes what the agent does, and the reward represents the evaluative signal associated with the interaction.
- State: the situation the agent is considering.
- Action: what the agent does.
- Reward: the evaluative signal associated with the interaction.
What the Framework Preserves
The framework is more informative than a description that lists only a final answer. It keeps the decision context, the agent's behavior, and the goal-related signal together. This allows the framework to represent cause and effect, because an action is considered in relation to what happens through the interaction. It also allows for uncertainty and nondeterminism, so a decision does not have to be treated as a perfectly predictable process. Finally, rewards provide explicit goals through an evaluative signal.
A Navigation Situation
Describing Navigation with Three Terms
Describe a simple situation in which an agent is navigating toward a target using states, actions, and rewards.
State: The agent's perceived state is "target ahead." This identifies the situation the agent is considering.
Action: The agent selects "move toward target." This identifies what the agent does.
Reward: The interaction produces the evaluative signal "progress toward target." This identifies how the outcome is described in relation to the goal.
Policy connection: The policy connects the perceived state "target ahead" with the action "move toward target."
The situation can be represented as state: "target ahead," action: "move toward target," and reward: "progress toward target."
This example is not a special policy format. Its labels are invented to make the mapping visible. The important structure is that a perceived state is connected to an action, while the reward supplies an evaluative signal associated with the interaction.
What do you think happens?
If the agent perceives the state "target ahead," what part of the framework identifies the behavior it performs?
Reveal answer
Answer: The policy
The policy provides the connection from the perceived state to the action. The state describes the situation, while the reward is the evaluative signal associated with the interaction.
Mistakes in Reading the Framework
Treating the policy as only the action
The action is what the agent does. The policy is the connection that maps a perceived state to an action.
Fix:
Describe the policy as the rule or mapping that connects "target ahead" with "move toward target."Leaving out the perceived state
An agent's behavior requires a connection from perception of the environment to an action.
Fix:
State the situation first, then identify the action connected to it by the policy.Confusing the reward with the action
The action describes what the agent does. The reward represents the evaluative signal associated with the interaction.
Fix:
Separate what the agent does from the signal that describes the desirability of the outcome.Assuming every policy must be a simple lookup table
A policy can be a simple function or lookup table, but producing the action can also involve extensive computation, such as a search process.
Fix:
Focus on the policy's role as a state-to-action mapping, not on one required representation.Describing only a final answer
The framework keeps the decision context, behavior, and goal-related signal together.
Fix:
Describe the situation with states, actions, and rewards and relate the action to what happens through the interaction.
Practice the Three-Part Description
Create your own simple reinforcement learning situation. Name one perceived state, one action connected to that state by a policy, and one reward that serves as an evaluative signal for the interaction.
Hints
- Begin with the situation the agent perceives.
- Separate what the agent does from the signal associated with the outcome.
- Explain how the policy connects the state to the action.
Explain the difference between a consistent policy and a stochastic policy. Use the same perceived state in both descriptions and state how the possible action selection differs.
Hints
- A consistent policy selects an action consistently for the state.
- A stochastic policy can involve randomness in its action selection.
- Both remain policies because both connect perceived states with actions.
The Framework in One Pass
- Start with the state the agent perceives. Apply the policy, which connects that state to an action. The agent and environment then continue their interaction, producing a reward as an evaluative signal. This compact framework keeps behavior, cause and effect, uncertainty and nondeterminism, and explicit goals in one description.
Key Takeaways
- A policy maps each perceived state to an action and therefore provides the connection that produces behavior.
- A policy can be a simple mapping or can involve extensive computation.
- A consistent policy selects an action consistently, while a stochastic policy can involve randomness in action selection.
- The reinforcement learning framework uses states, actions, and rewards to describe interaction between a learning agent and an environment.
- The framework preserves decision context, cause and effect, uncertainty and nondeterminism, and explicit goals.