Binary Classification
X names the collection of objects being studied, while Y names the collection of possible labels.
From Object to Prediction
A classification problem begins with two separate questions: What objects are we studying, and what labels could describe those objects? The collection of studied objects is the domain set, written as X. The collection of possible labels is the label set, written as Y. Binary classification uses a label set with two possible outcomes, while multiclass prediction has more than two possible target classes.
Domain Points and Feature Vectors
The domain set X contains the objects under study. An individual object in X is called a domain point. A domain point is usually represented by a vector of features. Each feature describes an aspect of that object for the learning problem. In the source example, a papaya is represented using features such as color and softness. These features describe the papaya; they are not labels.
Separating object, features, and label
Consider the papaya example. Identify which items belong to X, which items belong to Y, and which items describe a domain point.
Domain set: X is the set of all papayas, so every papaya, including one particular papaya, belongs on the domain side.
Feature representation: Color and softness are features used to represent a papaya as a domain point.
Label set: Y is {0, 1}, so 0 and 1 are the possible labels in this binary example.
The papaya is the object, color and softness are features, and 0 and 1 are possible labels.
Domain and Label Roles
The domain set X names the collection of objects being studied. The label set Y names the collection of possible labels. A domain object and a possible label play different roles: the object is what the problem examines, while the label is one of the allowed outcomes that can describe it.
When a domain point is considered together with a possible label, the object comes from X and the candidate outcome comes from Y. In the papaya example, a particular papaya remains a domain object even when it is considered together with 0 or with 1. The label does not become a feature, and the feature vector does not become a label.
Multiclass Prediction
A multiclass prediction problem has more than two possible target classes. Its goal is to learn a function that maps an input to one of k classes. Because a multiclass problem has several possible classes, it can be reduced to several binary classification problems, each asking a simpler yes-or-no question.
The reduction does not change the central distinction between domain objects and labels. The input is still a domain point represented by features, while the final prediction is selected from the problem's possible classes. What changes is the organization of the binary tasks used to reach that multiclass prediction.
One-Versus-All Training
The One-Versus-All Approach trains one binary classifier for each class. For a problem with k classes, it creates k binary training sets and k binary predictors. The classifier associated with class i learns to distinguish class i from every other class. Thus, each binary task gives one class its own contest against all remaining classes.
For a new instance, the trained predictors are considered together to construct a multiclass predictor. Each predictor represents one candidate class and distinguishes that class from all the remaining classes. The important point is the structure: there is one binary task per class, not one task for every pair of classes.
All-Pairs Training
The All-Pairs Approach trains a binary classifier for every pair of classes. The training set for a pair of classes i and j contains only examples whose original labels are i or j. Instead of asking one class to compete against all remaining classes, this approach gives every pair of classes a direct binary comparison.
The prediction process follows the same pairwise organization. The binary results from the relevant class comparisons are considered together to form a multiclass prediction. The source establishes that these results are combined differently from One-Versus-All; the essential distinction is that each result comes from a direct comparison between two classes.
Two Reduction Structures
| Aspect | One-Versus-All | All-Pairs |
|---|---|---|
| Binary training sets | One for each class | One for every pair of classes |
| Examples used by one task | The selected class versus all remaining classes | Only examples whose labels are the selected pair |
| Binary predictors for k classes | k predictors | A predictor for every class pair |
| Meaning of each binary task | One class competes against the rest | Two classes compete directly |
| Multiclass prediction | The class-specific predictors are considered together | The pairwise comparisons are considered together |
Common Classification Mistakes
Treating a domain object as a label
The papaya is an object in the domain set X, while the labels belong to Y.
Fix:
Place the papaya in X and reserve 0 and 1 for the possible labels in the source example.Treating features as labels
Color and softness are features used to represent a domain point.
Fix:
Use color and softness as feature components, and use 0 and 1 as the possible labels.Describing One-Versus-All as a pairwise method
One-Versus-All compares one class with every other class in its binary task.
Fix:
Remember that All-Pairs creates a separate task for each pair, while One-Versus-All creates one task for each class.Assuming both approaches use the same training examples in each binary task
An All-Pairs training set for classes i and j contains only examples whose original labels are i or j.
Fix:
Select only the examples belonging to the chosen pair for an All-Pairs task.
Check Your Understanding
A problem has three possible classes. Explain how One-Versus-All and All-Pairs would organize its binary training tasks. Then explain what kind of results would be considered together when making a multiclass prediction.
Hints
- For One-Versus-All, identify how many class-versus-rest tasks are created.
- For All-Pairs, list the kinds of class pairs that can be selected.
- Keep the training structure separate from the prediction process.
What do you think happens?
For a three-class problem, which approach creates a binary task whose training examples contain only classes 1 and 3?
Reveal answer
Answer: All-Pairs
The All-Pairs Approach creates a binary training set for every pair of classes, and the task for classes 1 and 3 contains only examples whose original labels are 1 or 3.
Key Takeaways
- X is the domain set of objects being studied; Y is the label set of possible labels.
- A domain point is usually represented by a feature vector, and features are not labels.
- In the source papaya example, X is the set of all papayas and Y is {0, 1}.
- A multiclass problem has more than two target classes and can be reduced to several binary problems.
- One-Versus-All trains one class-versus-rest classifier per class, while All-Pairs trains one classifier for every pair of classes.
- The two approaches differ both in which examples each binary task uses and in how the binary results are considered together for multiclass prediction.
Key Takeaways
- The domain set X contains the objects under study, while the label set Y contains possible labels.
- Feature vectors represent domain points; their features should not be confused with labels.
- Multiclass prediction can be reduced to several binary classification tasks.
- One-Versus-All creates one class-versus-rest task for each class.
- All-Pairs creates one task for every pair of classes and uses those pairwise results together for prediction.