Goal-Directed Learning
Reinforcement learning combines learning with decision-making in a computational setting.
Why Goals Change Learning
Many learning problems are not only about acquiring information. They also involve deciding what to do while working toward a goal. Reinforcement learning studies this kind of problem computationally: it combines learning with decision-making, and it focuses on an agent that improves through direct interaction with an environment.
The Reinforcement Learning Setting
A reinforcement-learning problem has an agent and an environment. The agent deals directly with the world around it, makes decisions, and learns through its interaction with that environment. Learning and decision-making are therefore considered together rather than treated as unrelated activities.
The agent is not learning in isolation. Its learning is connected to the decisions it must make while interacting with its environment and pursuing a goal over time.
Tracing Interaction and Feedback
Direct interaction is central because the agent learns while dealing with the environment itself. The agent does not merely receive a description of a task and remain separate from the situation. Instead, the computational problem is to improve decisions through interaction while working toward a goal.
An abstract interaction loop
Consider a learner that must make decisions while dealing directly with the world around it and working toward a goal that may take time to achieve.
1. The agent faces the environment: The learner is situated in an environment rather than learning separately from it.
2. The agent makes a decision: The learner must decide what to do as part of the problem.
3. Interaction produces information: The agent learns through its direct interaction with the environment.
4. Learning influences later decisions: The purpose of learning is connected to improving decisions while working toward the longer-term goal.
This is reinforcement learning as a computational setting: learning occurs through interaction, and the learning is tied to decision-making over time.
From Decisions to Long-Term Goals
A short-term decision and a long-term goal are not the same computational issue. Reinforcement learning focuses on situations in which an agent must learn from interaction in order to achieve goals extending over time. This means that a decision is understood not only as an isolated choice, but as part of an ongoing process of improving behavior toward a goal.
Three Computational Perspectives
Reinforcement learning is distinguished from approaches that depend on exemplary supervision or complete environment models. Its characteristic setting is learning through direct interaction while making decisions toward long-term goals. The contrast is about where the learning or decision-making support comes from.
| Approach | What the source emphasizes | What makes reinforcement learning distinct |
|---|---|---|
| Reinforcement learning | Learning through direct interaction while making decisions toward long-term goals | Learning and decision-making are combined in a computational setting |
| Methods depending on exemplary supervision | Dependence on exemplary supervision | They are not characterized here by the same direct-interaction setting |
| Methods depending on complete environment models | Dependence on a complete model of the environment | They are not characterized here by learning through direct interaction with the environment |
Common Misunderstandings
Treating reinforcement learning as learning without decisions
The source defines reinforcement learning as a combination of learning and decision-making.
Fix:
Always ask how learning is connected to the agent's decisions.Ignoring the environment
Direct interaction with the environment is central to the approach.
Fix:
Include the agent's interaction with its environment when describing the learning process.Reducing the goal to one immediate decision
The central challenge includes learning through interaction while working toward long-term goals.
Fix:
Connect current decisions with the longer-term goal the agent is trying to achieve.Defining reinforcement learning as requiring exemplary supervision or a complete environment model
The source distinguishes reinforcement learning from methods that depend on exemplary supervision or complete environment models.
Fix:
Describe reinforcement learning through its direct-interaction, learning-and-decision-making setting.
Check Your Understanding
A learner is described as improving its decisions while interacting directly with an environment and working toward a goal that may take time to achieve. Explain why this description fits reinforcement learning. Then state how the description would differ if it instead depended on exemplary supervision or a complete environment model.
Hints
- Start with the relationship among learning, decision-making, and interaction.
- Mention the role of the long-term goal.
- For the contrast, identify what each alternative depends on.
What do you think happens?
Before reading the answer, decide which description best captures reinforcement learning.
Reveal answer
Answer: Learning and decision-making are combined while an agent interacts directly with its environment.
The source describes reinforcement learning as a computational approach in which an agent learns through direct interaction with its environment while working toward long-term goals. It distinguishes this setting from methods that depend on exemplary supervision or complete environment models.
Key Takeaways
- Reinforcement learning combines learning with decision-making in a computational setting.
- An agent learns through direct interaction with its environment.
- The approach is distinguished from methods that depend on exemplary supervision or complete environment models.
- Its central challenge is improving decisions through interaction while pursuing goals that extend over time.
- Learning and decision-making must be considered together rather than as unrelated activities.
Key Takeaways
- Reinforcement learning is a computational approach that combines learning and decision-making.
- The agent learns through direct interaction with its environment.
- Its defining challenge is to improve decisions while pursuing long-term goals.
- It is distinguished from approaches that depend on exemplary supervision or complete environment models.