Reinforcement Learning vs Supervised Learning
Unsupervised learning is concerned with finding hidden structure in unlabeled data.
Three Learning Questions
A useful way to distinguish machine learning paradigms is to ask two questions: what information does the learner receive, and what is the learning process trying to produce? Supervised learning receives labeled examples from an external supervisor. Unsupervised learning looks for hidden structure in unlabeled data. Reinforcement learning learns through interaction and the agent's own experience while trying to maximize a reward signal.
Objectives Across Paradigms
The three paradigms pursue different results. Supervised learning uses externally supplied labeled examples and aims to generalize the correct responses represented by those examples to new situations. Unsupervised learning is concerned with finding hidden structure in unlabeled data. Reinforcement learning is concerned with maximizing a reward signal through the agent's interaction and experience. These objectives are different, so reinforcement learning is identified as its own paradigm alongside supervised and unsupervised learning.
| Paradigm | Information emphasized | Objective |
|---|---|---|
| Supervised learning | Externally supplied labeled examples | Generalize correct responses to new situations |
| Unsupervised learning | Unlabeled data | Find hidden structure |
| Reinforcement learning | Interaction and the agent's own experience | Maximize a reward signal |
Supervised Generalization
Supervised learning learns from externally supplied labeled examples. Each example provides a situation together with the correct response or label. The learner is expected to generalize from those known examples rather than simply memorize them, producing an appropriate response for situations that were not included in the training set.
A labeled-response setup
Imagine preparing examples that pair situations with the correct response, then presenting the learner with a situation that was not included in those examples.
Supply examples: An external supervisor provides situations together with their correct labels or actions.
Learn a relationship: The learner uses the labeled examples to learn how situations relate to correct responses.
Present a new situation: The learner encounters a situation outside the supplied training examples.
Generalize: The learner is expected to extend what it learned and produce an appropriate response for the new situation.
This is the supervised-learning pattern: externally supplied labels support generalization to situations not included in the training set.
Interaction and Experience
Reinforcement learning starts with an agent that interacts with an environment. At a high level, the agent encounters a situation, responds, receives feedback through what happens, and uses that experience as part of further learning. It does not depend on a complete collection of correct responses for every situation in advance. This matters when the agent enters situations for which representative desired responses may be difficult to prepare.
An agent learning by interaction
Imagine an agent placed in situations where it must respond, without a complete collection of correct responses prepared in advance.
Encounter: The agent encounters a situation in its environment.
Respond: The agent chooses a response rather than receiving a supplied correct action for every possible situation.
Experience: The agent uses what happens, including the reward signal, as part of its own experience.
Continue learning: That experience contributes to later behavior as the agent continues interacting with situations.
This is the reinforcement-learning pattern: knowledge develops through interaction and experience while the learner is concerned with maximizing a reward signal.
Information at Learning Time
The information available during learning is a central difference. A supervised learner is shown examples that already specify what the correct response should be. An agent learning through reinforcement does not depend on having a complete collection of correct responses for every situation in advance. Instead, it gains knowledge while interacting with situations and using its own experience.
| Question | Supervised learning | Reinforcement learning |
|---|---|---|
| Where does learning information come from? | An external supervisor supplies labeled examples. | The agent learns through interaction and its own experience. |
| Is the correct response supplied for every situation in advance? | The examples specify correct responses for the supplied cases. | The agent does not depend on a complete collection of correct responses for every situation in advance. |
| What must happen beyond the known cases? | The learner generalizes to situations not included in the training set. | The agent continues learning while it encounters and responds to situations. |
Why Unsupervised Is Different
Reinforcement learning can initially look like unsupervised learning because both may lack examples that explicitly show the correct action. That first impression focuses on what reinforcement learning does not use. The more complete classification asks what the process is trying to produce. Unsupervised learning seeks hidden structure in unlabeled data. Reinforcement learning seeks to maximize a reward signal through interaction. Since their objectives differ, reinforcement learning is not a subtype of unsupervised learning.
When classifying a learning method, do not stop at the question, “Are correct examples present?” Also ask what the learner is trying to produce. Hidden structure points toward unsupervised learning; generalization from supplied labels points toward supervised learning; maximizing a reward signal through interaction points toward reinforcement learning.
Limits of Labeled Examples
Labeled examples are useful, but they do not by themselves solve every learning problem. In an interactive problem, the system may need to act across many situations. Preparing examples of desired behavior for all of those situations can be impractical, especially when the examples must be correct and representative of the situations the agent will encounter. Reinforcement learning uses a different starting point: the agent learns while interacting with situations, including situations for which a prewritten collection of representative desired responses may not be available.
Common Classification Mistakes
Calling reinforcement learning unsupervised simply because it does not use examples that explicitly show the correct action.
It classifies the method only by what it lacks and ignores the objective.
Fix:
Check whether the process seeks hidden structure or seeks to maximize a reward signal through interaction.Treating supervised learning as memorizing the training set.
Supervised learning is expected to generalize correct responses to situations not included in the training set.
Fix:
Include generalization to new situations in the description of supervised learning.Describing reinforcement learning as receiving the correct response for every situation.
Reinforcement learning does not depend on a complete collection of correct responses for every situation in advance.
Fix:
Describe the agent as learning through interaction, feedback, and its own experience.Ignoring the objective when comparing paradigms.
The paradigms pursue different results: correct-response generalization, hidden structure, and reward maximization.
Fix:
Compare both the information available and what the learner is trying to produce.
Classification Practice
For each description, identify the best-fitting paradigm and explain your reason. A learner receives externally supplied situations paired with correct responses. A learner searches unlabeled data for hidden structure. An agent encounters situations, chooses responses, and uses a reward signal and its own experience for further learning.
Hints
- Look first at the information supplied to the learner.
- Then identify whether the objective is correct-response generalization, hidden structure, or reward maximization.
- Remember that the absence of correct-action examples does not by itself identify unsupervised learning.
Practice classification
Classify the three descriptions from the practice prompt.
Description one: Externally supplied situations paired with correct responses describe supervised learning because the learner uses labeled examples.
Description two: Searching unlabeled data for hidden structure describes unsupervised learning.
Description three: An agent choosing responses and learning from a reward signal and its own experience describes reinforcement learning.
The classifications are supervised learning, unsupervised learning, and reinforcement learning, respectively.
Key Takeaways
- Supervised learning learns from externally supplied labeled examples and generalizes correct responses to new situations.
- Unsupervised learning seeks hidden structure in unlabeled data.
- Reinforcement learning learns through interaction and the agent's own experience while trying to maximize a reward signal.
- Reinforcement learning is not a subtype of unsupervised learning because its objective differs: it is concerned with reward rather than hidden structure.
- For reliable classification, consider both the information available to the learner and the result the learning process is trying to produce.
Key Takeaways
- Supervised learning uses externally supplied labeled examples and generalizes from them to new situations.
- Unsupervised learning finds hidden structure in unlabeled data.
- Reinforcement learning learns through interaction and the agent's own experience to maximize a reward signal.
- The lack of correct-action examples does not make reinforcement learning unsupervised.
- Reinforcement learning is a distinct paradigm alongside supervised and unsupervised learning.