Concepts / Reinforcement Learning vs Supervised Learning

Reinforcement Learning vs Supervised Learning

Unsupervised learning is concerned with finding hidden structure in unlabeled data.

  • Programming

Three Learning Questions

A useful way to distinguish machine learning paradigms is to ask two questions: what information does the learner receive, and what is the learning process trying to produce? Supervised learning receives labeled examples from an external supervisor. Unsupervised learning looks for hidden structure in unlabeled data. Reinforcement learning learns through interaction and the agent's own experience while trying to maximize a reward signal.

Objectives Across Paradigms

The three paradigms pursue different results. Supervised learning uses externally supplied labeled examples and aims to generalize the correct responses represented by those examples to new situations. Unsupervised learning is concerned with finding hidden structure in unlabeled data. Reinforcement learning is concerned with maximizing a reward signal through the agent's interaction and experience. These objectives are different, so reinforcement learning is identified as its own paradigm alongside supervised and unsupervised learning.

Supervised learningCorrect responsesUnsupervised learningHidden structureReinforcementlearningReward signal
How do the goals differ between predicting supplied labels, finding structure in unlabeled data, and maximizing reward through actions?
ParadigmInformation emphasizedObjective
Supervised learningExternally supplied labeled examplesGeneralize correct responses to new situations
Unsupervised learningUnlabeled dataFind hidden structure
Reinforcement learningInteraction and the agent's own experienceMaximize a reward signal

Supervised Generalization

Supervised learning learns from externally supplied labeled examples. Each example provides a situation together with the correct response or label. The learner is expected to generalize from those known examples rather than simply memorize them, producing an appropriate response for situations that were not included in the training set.

suppliesgeneralizes toproducesLabeled examplesKnown responsesLearningGeneralizationNew situationNot in training setResponseCorrect action or label
How does a supervised learner transform known labeled examples into a response for an input not included in training?

A labeled-response setup

Imagine preparing examples that pair situations with the correct response, then presenting the learner with a situation that was not included in those examples.

Supply examples: An external supervisor provides situations together with their correct labels or actions.

Learn a relationship: The learner uses the labeled examples to learn how situations relate to correct responses.

Present a new situation: The learner encounters a situation outside the supplied training examples.

Generalize: The learner is expected to extend what it learned and produce an appropriate response for the new situation.

This is the supervised-learning pattern: externally supplied labels support generalization to situations not included in the training set.

Interaction and Experience

Reinforcement learning starts with an agent that interacts with an environment. At a high level, the agent encounters a situation, responds, receives feedback through what happens, and uses that experience as part of further learning. It does not depend on a complete collection of correct responses for every situation in advance. This matters when the agent enters situations for which representative desired responses may be difficult to prepare.

chooseleads tobecomesinfluences next responseSituationAgent encountersActionAgent respondsFeedbackWhat happensExperienceFurther learning
How does an agent observe a situation, choose an action, receive feedback, and use the result to influence what it does next?

An agent learning by interaction

Imagine an agent placed in situations where it must respond, without a complete collection of correct responses prepared in advance.

Encounter: The agent encounters a situation in its environment.

Respond: The agent chooses a response rather than receiving a supplied correct action for every possible situation.

Experience: The agent uses what happens, including the reward signal, as part of its own experience.

Continue learning: That experience contributes to later behavior as the agent continues interacting with situations.

This is the reinforcement-learning pattern: knowledge develops through interaction and experience while the learner is concerned with maximizing a reward signal.

Information at Learning Time

The information available during learning is a central difference. A supervised learner is shown examples that already specify what the correct response should be. An agent learning through reinforcement does not depend on having a complete collection of correct responses for every situation in advance. Instead, it gains knowledge while interacting with situations and using its own experience.

suppliesleads toproducescontributes toLabeled examplesCorrect response suppliedSituationAgent encountersSupervised learnerGeneralizesActionAgent choosesReward signalFeedbackSubsequentsituationFurther experience
What information does a supervised learner receive from labeled examples, and what information does a reinforcement-learning agent receive through interaction?
QuestionSupervised learningReinforcement learning
Where does learning information come from?An external supervisor supplies labeled examples.The agent learns through interaction and its own experience.
Is the correct response supplied for every situation in advance?The examples specify correct responses for the supplied cases.The agent does not depend on a complete collection of correct responses for every situation in advance.
What must happen beyond the known cases?The learner generalizes to situations not included in the training set.The agent continues learning while it encounters and responds to situations.

Why Unsupervised Is Different

Reinforcement learning can initially look like unsupervised learning because both may lack examples that explicitly show the correct action. That first impression focuses on what reinforcement learning does not use. The more complete classification asks what the process is trying to produce. Unsupervised learning seeks hidden structure in unlabeled data. Reinforcement learning seeks to maximize a reward signal through interaction. Since their objectives differ, reinforcement learning is not a subtype of unsupervised learning.

revealssupports maximizingUnlabeled dataNo supplied labelsInteractionAgent experienceHidden structureLearning objectiveReward signalLearning objective
What distinguishes finding hidden structure in unlabeled data from learning which actions lead to higher rewards?

When classifying a learning method, do not stop at the question, “Are correct examples present?” Also ask what the learner is trying to produce. Hidden structure points toward unsupervised learning; generalization from supplied labels points toward supervised learning; maximizing a reward signal through interaction points toward reinforcement learning.

Limits of Labeled Examples

Labeled examples are useful, but they do not by themselves solve every learning problem. In an interactive problem, the system may need to act across many situations. Preparing examples of desired behavior for all of those situations can be impractical, especially when the examples must be correct and representative of the situations the agent will encounter. Reinforcement learning uses a different starting point: the agent learns while interacting with situations, including situations for which a prewritten collection of representative desired responses may not be available.

maps torequiresencountersSupplied situationLabeled responseEncounteredsituationAgent experienceCorrect responseExternal labelChosen actionAgent respondsConsequenceReward signal
What changes when a learner must choose actions and encounter consequences rather than simply map fixed inputs to supplied labels?

Common Classification Mistakes

  • Calling reinforcement learning unsupervised simply because it does not use examples that explicitly show the correct action.

    It classifies the method only by what it lacks and ignores the objective.

    Fix: Check whether the process seeks hidden structure or seeks to maximize a reward signal through interaction.

  • Treating supervised learning as memorizing the training set.

    Supervised learning is expected to generalize correct responses to situations not included in the training set.

    Fix: Include generalization to new situations in the description of supervised learning.

  • Describing reinforcement learning as receiving the correct response for every situation.

    Reinforcement learning does not depend on a complete collection of correct responses for every situation in advance.

    Fix: Describe the agent as learning through interaction, feedback, and its own experience.

  • Ignoring the objective when comparing paradigms.

    The paradigms pursue different results: correct-response generalization, hidden structure, and reward maximization.

    Fix: Compare both the information available and what the learner is trying to produce.

Classification Practice

EASY

For each description, identify the best-fitting paradigm and explain your reason. A learner receives externally supplied situations paired with correct responses. A learner searches unlabeled data for hidden structure. An agent encounters situations, chooses responses, and uses a reward signal and its own experience for further learning.

Hints
  • Look first at the information supplied to the learner.
  • Then identify whether the objective is correct-response generalization, hidden structure, or reward maximization.
  • Remember that the absence of correct-action examples does not by itself identify unsupervised learning.

Practice classification

Classify the three descriptions from the practice prompt.

Description one: Externally supplied situations paired with correct responses describe supervised learning because the learner uses labeled examples.

Description two: Searching unlabeled data for hidden structure describes unsupervised learning.

Description three: An agent choosing responses and learning from a reward signal and its own experience describes reinforcement learning.

The classifications are supervised learning, unsupervised learning, and reinforcement learning, respectively.

Key Takeaways

  1. Supervised learning learns from externally supplied labeled examples and generalizes correct responses to new situations.
  2. Unsupervised learning seeks hidden structure in unlabeled data.
  3. Reinforcement learning learns through interaction and the agent's own experience while trying to maximize a reward signal.
  4. Reinforcement learning is not a subtype of unsupervised learning because its objective differs: it is concerned with reward rather than hidden structure.
  5. For reliable classification, consider both the information available to the learner and the result the learning process is trying to produce.

Key Takeaways

  • Supervised learning uses externally supplied labeled examples and generalizes from them to new situations.
  • Unsupervised learning finds hidden structure in unlabeled data.
  • Reinforcement learning learns through interaction and the agent's own experience to maximize a reward signal.
  • The lack of correct-action examples does not make reinforcement learning unsupervised.
  • Reinforcement learning is a distinct paradigm alongside supervised and unsupervised learning.