Supervised Learning
Unsupervised learning finds useful transformations of input data without target values.
A Dataset Without Answers
Machine-learning problems differ first by what the data contains. If each input example is paired with the answer the system is expected to learn, the problem is supervised learning. If the dataset contains observations but no target value for each observation, the setting is unsupervised learning. This difference changes what the learner is trying to discover.
The central question is not simply whether data is available. It is whether the intended answer is available alongside each input example.
Investigating Unlabeled Data
Unsupervised learning finds useful transformations of input data without target values. Imagine receiving observations with no column that tells you the correct answer for each one. You can still investigate the dataset by looking for a more useful representation, making the data easier to view, reducing its size, removing noise, or examining relationships among its measurements.
Two well-known categories are dimensionality reduction and clustering. They begin with input data that has no targets, but they ask different questions. Dimensionality reduction changes how each example is represented. Clustering looks across examples to discover which data points appear to belong together.
| Category | Question | Output |
|---|---|---|
| Dimensionality reduction | Can each observation be represented with fewer dimensions? | A representation with fewer dimensions |
| Clustering | Which observations appear to belong together? | Groups of similar data points |
The distinction is determined by what the learner produces from the unlabeled input data.
Why Explore Before Prediction
Unsupervised learning can be useful before attempting a supervised-learning problem because it helps analysts investigate the available data without needing target values. A transformed representation may make the data easier to visualize or store. Other unsupervised uses include data compression, data denoising, and understanding correlations present in the measurements.
Source example: A dataset records many measurements for each observation. If a learner needs a representation with fewer dimensions so the data is easier to visualize or store, the task is dimensionality reduction. If the goal changes to discovering which observations appear to belong together, the task is clustering.
The Input-to-Target Mapping
Supervised learning connects inputs with known targets through a collection of examples. The targets are also called annotations. The learner uses these paired examples to discover a mapping from input data to known targets.
Recognizing the Learning Setup
A collection contains inputs, and every input is accompanied by the answer the system is expected to learn. Identify the learning setting and describe the role of the data.
Inspect the examples: Each input is paired with an answer rather than appearing alone.
Name the target: The paired answer is the known target, also called an annotation.
Describe the learned result: The learner uses the examples to discover a mapping from inputs to targets.
Apply the mapping: When a new input is presented, the mapping is used to produce a prediction.
This is supervised learning because the inputs and their known targets are supplied together as examples.
What do you think happens?
A dataset contains observations but no target value for each observation. Is it enough to call the problem supervised learning?
Reveal answer
Answer: No, because supervised learning requires known targets paired with inputs.
Supervised learning begins with examples that connect inputs and known targets. A dataset without target values is the setting for unsupervised learning.
Reading the Target Shape
The target does not have to be a single category. The structure of the target determines the supervised-learning task. Common task forms include classification and regression, while more specialized targets can be sequences, syntax trees, bounding boxes, or pixel-level masks.
| Task | Target form | What the target describes |
|---|---|---|
| Classification | Category | Which category is associated with the input |
| Regression | Numeric target | A value associated with the input |
| Sequence generation | Ordered sequence | A complete ordered output |
| Syntax tree prediction | Nested structure | The structure of a sentence |
| Object detection | Objects and bounding boxes | Which objects are present and where they are located |
| Image segmentation | Pixel-level mask | The pixels belonging to a specific object |
These tasks share the supervised-learning pattern but differ in the structure of their known targets.
Specialized Targets
Sequence generation predicts an ordered output such as a caption. One useful interpretation is to treat generation as repeated classification: the system repeatedly predicts the next word or token. The final target is still a sequence, even though the prediction can be organized as a series of one-element-at-a-time choices.
Object detection identifies objects with bounding boxes. It can be expressed as classification, such as classifying candidate boxes, or as a joint classification and regression problem in which box coordinates are predicted through vector regression. Image segmentation describes the target more precisely: its target is a pixel-level mask marking a specific object.
Source example: A supervised dataset may pair an image with a target that identifies objects and places bounding boxes around them. A different image task may pair the image with a mask that marks the exact pixels belonging to an object. Both are supervised because each input is accompanied by a known target.
The same supervised-learning framework supports many applications. What changes is the target structure: a category, a numeric value, an ordered sequence, a nested tree, object boxes, or a pixel mask.
Mistakes About Labels and Outputs
Treating every learning problem as supervised because it has input data.
Supervised learning requires inputs paired with known targets. Input data without targets is the setting for unsupervised learning.
Fix:
Check whether the intended answer or annotation is supplied alongside each example.Confusing dimensionality reduction with clustering.
Dimensionality reduction seeks a representation with fewer dimensions, while clustering seeks groups among data points.
Fix:
Ask whether the output changes each example's representation or assigns observations to groups.Assuming supervised learning always predicts one category.
Supervised targets can be sequences, nested structures, boxes, or masks as well as categories and numeric values.
Fix:
Identify the target form and verify that it is known for each input example.Treating object detection and image segmentation as the same target.
Object detection uses bounding boxes, while image segmentation uses a pixel-level mask.
Fix:
Distinguish location boxes from exact pixel markings.
Check the Learning Setup
For each description, identify whether it is unsupervised learning or supervised learning. If it is supervised, name the target form. 1. Observations are examined to discover groups of data points, but no answers are supplied. 2. Each input is paired with a category. 3. Each image is paired with a pixel-level mask for a specific object. 4. A system learns to produce an ordered caption from paired examples. 5. A representation with fewer dimensions is sought so the data is easier to visualize or store.
Hints
- First check whether known targets are supplied.
- For supervised tasks, focus on the structure of the target.
- Groups and fewer dimensions are different unsupervised outputs.
Practice Answers
Classify the five learning descriptions by their data setup and target form.
1: Unsupervised learning: the goal is to discover groups without target values.
2: Supervised learning with a category target, so it is classification.
3: Supervised learning with a pixel-level mask target, so it is image segmentation.
4: Supervised learning with an ordered sequence target, so it is sequence generation.
5: Unsupervised learning through dimensionality reduction, because the desired output has fewer dimensions.
The first and fifth descriptions are unsupervised. The middle three are supervised, and their target forms are a category, a pixel mask, and an ordered sequence.
Key Takeaways
- Unsupervised learning works with input data without target values and can reveal useful structure or transformations.
- Dimensionality reduction seeks fewer dimensions, while clustering seeks groups among data points.
- Unsupervised learning can support visualization, compression, denoising, and understanding correlations before a supervised-learning problem.
- Supervised learning connects inputs with known targets through examples and uses the learned input-to-target mapping to produce predictions.
- The target structure determines the task: categories, numeric values, sequences, syntax trees, bounding boxes, and pixel masks all fit the supervised-learning framework.
Key Takeaways
- Supervised learning begins with inputs paired with known targets; unsupervised learning begins with input data without targets.
- Dimensionality reduction changes the representation of each example, while clustering discovers groups among examples.
- Unsupervised methods can help visualize, compress, denoise, and understand correlations in data before prediction is attempted.
- Supervised-learning targets may be categories, numbers, sequences, syntax trees, bounding boxes, or pixel-level masks.
- Sequence generation, object detection, syntax tree prediction, and image segmentation are supervised because their target outputs are known for the training examples.