Branches of Machine Learning
Supervised learning connects inputs with known targets through a collection of examples.
From Examples to Predictions
Suppose every input in a collection comes with the answer a system is expected to learn. That pairing is the starting point for supervised learning. The system studies examples that connect inputs with known targets, discovers a mapping between them, and then uses that mapping to produce a prediction for a new input.
The defining question is whether the intended answer is available alongside each input example. Known targets are also called annotations.
Tracing the Learning Process
- Begin with a collection of inputs and their known targets.
- Use the paired examples to discover a mapping from input data to targets.
- Represent that mapping in a model.
- Present a new input to the model.
- Use the mapping to produce a prediction.
The role of the data changes during supervised learning. At the beginning, the data consists of labeled examples: each input is accompanied by a known target. After learning, the result is a reusable input-to-target relationship. The model can apply that relationship to a new input, even though the new input was not one of the original examples.
A Captioning Example
Consider a collection of images in which each image is paired with a caption. What makes this a supervised-learning setup?
Identify the input: Each image is an input example.
Identify the target: The caption paired with an image is its known target.
Learn the relationship: The system uses the paired image-and-caption examples to learn a mapping from an image to a caption.
Predict for a new input: When a new image is presented, the learned mapping is used to produce a caption.
The setup is supervised because every training input is accompanied by the answer the system is expected to learn.
Target Shapes and Task Types
| Task type | Target structure | What the system learns to produce |
|---|---|---|
| Classification | A category | A category associated with the input |
| Regression | A value or vector of values | A numerical target |
| Sequence generation | An ordered sequence | A sequence such as a caption |
| Syntax tree prediction | A nested tree structure | The hierarchical structure of a sentence |
| Object detection | Object classes and bounding boxes | Identified objects and their locations |
| Image segmentation | A pixel-level mask | The pixels belonging to a specific object |
Classification and regression are common supervised-learning tasks, but the same input-with-known-target structure supports more specialized tasks. A sequence target preserves order. A syntax-tree target represents nested structure. An image target may describe object locations or mark exact pixels.
Sequence generation can be viewed as repeated classification. Instead of treating the whole output as one indivisible choice, the system repeatedly predicts the next word or token. The final output remains an ordered sequence, but it is produced through a series of choices made one element at a time.
Structured Image Targets
Object detection has a richer target than ordinary classification. Its target identifies objects and includes bounding boxes around them. It can be expressed as classification, by classifying candidate boxes, or as a joint classification-and-regression problem in which box coordinates are predicted through vector regression.
Syntax tree prediction is supervised learning when each input sentence is paired with the hierarchical syntax tree that the system is expected to learn. The target is not merely one label; it is a nested structure describing the sentence.
Image segmentation goes further than object detection. Detection identifies objects and surrounds them with bounding boxes. Segmentation marks the exact pixels belonging to a specific object, so its target is a pixel-level mask.
Common Misreadings
Assuming supervised learning must predict a single category.
The target can have many structures. The defining feature is the pairing of each input with a known target.
Fix:
Inspect the form of the target. It may be a category, value, sequence, tree, box, or mask.Confusing object detection with image segmentation.
A bounding box describes an object's location with a box, whereas segmentation marks the exact pixels belonging to a specific object.
Fix:
Use detection for identified objects and bounding boxes; use segmentation for pixel-level masks.Treating sequence generation as one unstructured output.
A sequence is an ordered target, and sequence generation can be understood as repeated prediction of the next word or token.
Fix:
Track both the target's ordered structure and the one-element-at-a-time interpretation.Calling data supervised merely because inputs exist.
Supervised learning begins with inputs paired with known targets or annotations.
Fix:
Ask whether a known target is available alongside every example.
Practice the Target Test
For each scenario, identify the target structure and the supervised-learning task it represents: a sentence paired with a nested syntax tree; an image paired with object classes and bounding boxes; an image paired with a pixel-level object mask; and an input paired with an ordered caption.
Hints
- First identify what answer is paired with the input.
- Then classify the answer as a category, value, sequence, tree, box-based description, or mask.
- Remember that a caption is ordered, while a syntax tree is hierarchical.
Applying the Target Test
Classify four input-target pairings by their target structure.
Sentence and nested structure: The target is a syntax tree, so this is syntax tree prediction.
Image and object locations: The target contains object classes and bounding boxes, so this is object detection.
Image and exact object pixels: The target is a pixel-level mask, so this is image segmentation.
Input and ordered caption: The target is a sequence, so this is sequence generation.
All four are supervised-learning tasks because each input is paired with a known target. Their task types differ because their targets have different structures.
Key Takeaways
- Supervised learning connects inputs with known targets through a collection of examples.
- The learned result is a reusable mapping from inputs to targets that can produce predictions for new inputs.
- The target's structure determines the task, including classification, regression, sequence generation, syntax tree prediction, object detection, and image segmentation.
- Sequence generation can be viewed as repeated classification of the next word or token.
- Object detection predicts object identities and bounding boxes, while image segmentation predicts a pixel-level mask.
Key Takeaways
- Supervised learning starts with inputs paired with known targets, also called annotations.
- Learning changes those labeled examples into a reusable input-to-target mapping.
- Classification, regression, sequences, trees, boxes, and masks are different target structures supported by supervised learning.
- Sequence generation predicts ordered elements, object detection predicts classes and bounding boxes, syntax prediction produces hierarchical trees, and segmentation marks object pixels.