Goals and Uncertainty in Artificial Intelligence
The reinforcement learning framework describes interaction between a learning agent and an environment.
The Decision Problem
When an intelligent system must act, three questions matter: What situation is it in? What can it do? What result tells it whether the outcome was desirable? The reinforcement learning framework organizes these questions around a learning agent and its environment. Instead of describing an artificial intelligence problem as one large, vague task, it represents the interaction using states, actions, and rewards.
The framework is not just a description of a final answer. It keeps the situation, the agent's behavior, and the goal-related signal together.
Interaction Loop
Reinforcement learning describes a continuing interaction between two participants: a learning agent and an environment. The environment provides a state, the agent selects an action, and the interaction produces a reward. State, action, and reward are the vocabulary used to describe what passes between the participants and what the interaction means.
One Step Under Uncertainty
The framework connects an agent's action with what happens through the interaction. This gives it a way to represent cause and effect: an action is considered in relation to its consequences. The framework also allows uncertainty and nondeterminism. After an action, the result does not have to be treated as perfectly predictable; the interaction may be described as leading to different possible next situations and reward signals.
Uncertainty does not remove the action–outcome relationship. It means the framework can preserve the possibility of different outcomes instead of assuming that every decision has one perfectly predictable result.
Goals Through Rewards
Rewards provide the evaluative signal associated with the interaction. They allow the description to include an explicit goal: the outcome can be considered in relation to whether it is desirable. In this way, the framework does not merely record what the agent did or which situation followed. It also records the goal-related signal connected with that interaction.
A useful reading order is: first identify the situation, then identify what the agent does, and finally identify the signal that describes the desirability of the result. This sequence connects the agent's situation, its chosen action, and the associated reward.
Navigation Mapping
Consider a generated example in which an agent is navigating a space. The framework does not describe the entire navigation problem as one undivided task. Instead, describe one interaction by naming the agent's current situation, the choice available to it, and the signal associated with what happens next.
Mapping a Navigation Situation
An agent is at a location in a space and must choose between two available movements. Represent this situation with states, actions, and rewards.
State: Represent the agent's current location and situation as the state being considered.
Action: Represent the movement selected by the agent as the action.
Outcome: Represent the resulting situation as a possible next state. Because the framework allows uncertainty and nondeterminism, the action need not be treated as producing only one perfectly predictable result.
Reward: Represent the signal associated with the interaction as the reward. This signal indicates how the outcome relates to the goal being represented.
The navigation interaction can be described as current state, chosen action, possible next state, and associated reward. The description preserves the agent, the environment, cause and effect, uncertainty, and an explicit goal-related signal.
Common Misreadings
Treating the framework as a list of final answers
The framework keeps the decision context, the agent's behavior, and the goal-related signal together.
Fix:
Describe the state, the action, and the associated reward.Confusing the participants with the vocabulary
The agent and environment are the two interacting participants. State, action, and reward describe their interaction.
Fix:
Name who participates and separately name the situation, behavior, and evaluative signal.Assuming perfect predictability
The framework allows uncertainty and nondeterminism.
Fix:
Allow the description to include different possible outcomes and associated rewards.Leaving out the goal-related signal
Rewards represent the evaluative signal and make explicit goals part of the description.
Fix:
Include the reward associated with the interaction.
When analyzing a new situation, use the same three-question checklist: What situation is the agent in? What action can it take? What reward indicates the desirability of the resulting interaction?
Apply the Framework
Describe a simple artificial intelligence situation of your own using the reinforcement learning framework. Identify the learning agent, the environment, one possible state, one possible action, and the associated reward. Then explain where cause and effect, uncertainty or nondeterminism, and the explicit goal appear in your description.
Hints
- Start with the agent's situation rather than with the final result.
- Name the action separately from the state.
- State what the reward tells the agent about the desirability of the outcome.
- If the outcome is not perfectly predictable, describe more than one possible next situation.
Framework Summary
- Reinforcement learning represents interaction between a learning agent and an environment.
- The three central components are states, actions, and rewards.
- A state describes the situation, an action describes what the agent does, and a reward represents the associated evaluative signal.
- The framework connects actions with outcomes, while allowing uncertainty and nondeterminism.
- Rewards make explicit goals part of the description.
Key Takeaways
- A reinforcement learning framework describes an interacting learning agent and environment.
- States, actions, and rewards provide the framework's core vocabulary.
- The framework links an agent's action to possible outcomes and associated reward signals.
- Uncertainty and nondeterminism can be represented instead of assuming perfect predictability.
- Rewards preserve explicit goals by indicating how the interaction's outcome is evaluated.