Concepts / Binary Classification

Binary Classification

X names the collection of objects being studied, while Y names the collection of possible labels.

  • Programming

From Object to Prediction

A classification problem begins with two separate questions: What objects are we studying, and what labels could describe those objects? The collection of studied objects is the domain set, written as X. The collection of possible labels is the label set, written as Y. Binary classification uses a label set with two possible outcomes, while multiclass prediction has more than two possible target classes.

Domain Points and Feature Vectors

The domain set X contains the objects under study. An individual object in X is called a domain point. A domain point is usually represented by a vector of features. Each feature describes an aspect of that object for the learning problem. In the source example, a papaya is represented using features such as color and softness. These features describe the papaya; they are not labels.

represented bycontainscontainsPapayadomain point in XFeature vectorrepresentationColorfeature positionSoftnessfeature position
How does a domain object become a vector of feature values, and what does each vector position represent?

Separating object, features, and label

Consider the papaya example. Identify which items belong to X, which items belong to Y, and which items describe a domain point.

Domain set: X is the set of all papayas, so every papaya, including one particular papaya, belongs on the domain side.

Feature representation: Color and softness are features used to represent a papaya as a domain point.

Label set: Y is {0, 1}, so 0 and 1 are the possible labels in this binary example.

The papaya is the object, color and softness are features, and 0 and 1 are possible labels.

Domain and Label Roles

The domain set X names the collection of objects being studied. The label set Y names the collection of possible labels. A domain object and a possible label play different roles: the object is what the problem examines, while the label is one of the allowed outcomes that can describe it.

containscontainscontainsXall papayasPapayaone domain object0possible labelY{0, 1}1possible label
What is contained in the domain set X versus the label set Y, and why are domain objects different from the labels that can be assigned to them?

When a domain point is considered together with a possible label, the object comes from X and the candidate outcome comes from Y. In the papaya example, a particular papaya remains a domain object even when it is considered together with 0 or with 1. The label does not become a feature, and the feature vector does not become a label.

combined withcombined withxdomain object(x, y)object considered withlabelypossible label
How is a domain object considered together with a possible label, such as an object x paired with a label y?

Multiclass Prediction

A multiclass prediction problem has more than two possible target classes. Its goal is to learn a function that maps an input to one of k classes. Because a multiclass problem has several possible classes, it can be reduced to several binary classification problems, each asking a simpler yes-or-no question.

The reduction does not change the central distinction between domain objects and labels. The input is still a domain point represented by features, while the final prediction is selected from the problem's possible classes. What changes is the organization of the binary tasks used to reach that multiclass prediction.

One-Versus-All Training

The One-Versus-All Approach trains one binary classifier for each class. For a problem with k classes, it creates k binary training sets and k binary predictors. The classifier associated with class i learns to distinguish class i from every other class. Thus, each binary task gives one class its own contest against all remaining classes.

separate class 1separate class 2separate class ktraintraintrainMulticlass setclasses 1 through kClass 1 versus restbinary training setk predictorsone for each classClass 2 versus restbinary training setClass k versus restbinary training set
How is one multiclass training set transformed into several binary training sets, each separating one class from all remaining classes?

For a new instance, the trained predictors are considered together to construct a multiclass predictor. Each predictor represents one candidate class and distinguishes that class from all the remaining classes. The important point is the structure: there is one binary task per class, not one task for every pair of classes.

All-Pairs Training

The All-Pairs Approach trains a binary classifier for every pair of classes. The training set for a pair of classes i and j contains only examples whose original labels are i or j. Instead of asking one class to compete against all remaining classes, this approach gives every pair of classes a direct binary comparison.

select pairselect pairselect pairbinary resultbinary resultbinary resultcontribute toClassesmulticlass problemClasses i and jexamples labeled i or jPairwise resultscombined for predictionMulticlass predictionone classClasses i and kexamples labeled i or kClasses j and kexamples labeled j or k
How are binary training sets created for every pair of classes, and how do their comparisons contribute to a multiclass prediction?

The prediction process follows the same pairwise organization. The binary results from the relevant class comparisons are considered together to form a multiclass prediction. The source establishes that these results are combined differently from One-Versus-All; the essential distinction is that each result comes from a direct comparison between two classes.

Two Reduction Structures

AspectOne-Versus-AllAll-Pairs
Binary training setsOne for each classOne for every pair of classes
Examples used by one taskThe selected class versus all remaining classesOnly examples whose labels are the selected pair
Binary predictors for k classesk predictorsA predictor for every class pair
Meaning of each binary taskOne class competes against the restTwo classes compete directly
Multiclass predictionThe class-specific predictors are considered togetherThe pairwise comparisons are considered together
createsproducescreatesproducesOne-Versus-Alltrainingone binary task per classk predictorsclass versus restClass-specificresultsconsidered togetherAll-Pairs trainingone binary task per pairPairwise predictorsone for every class pairPairwise resultsconsidered together
What differs between the number and structure of binary models trained and the sequence of decisions made during prediction in the two approaches?

Common Classification Mistakes

  • Treating a domain object as a label

    The papaya is an object in the domain set X, while the labels belong to Y.

    Fix: Place the papaya in X and reserve 0 and 1 for the possible labels in the source example.

  • Treating features as labels

    Color and softness are features used to represent a domain point.

    Fix: Use color and softness as feature components, and use 0 and 1 as the possible labels.

  • Describing One-Versus-All as a pairwise method

    One-Versus-All compares one class with every other class in its binary task.

    Fix: Remember that All-Pairs creates a separate task for each pair, while One-Versus-All creates one task for each class.

  • Assuming both approaches use the same training examples in each binary task

    An All-Pairs training set for classes i and j contains only examples whose original labels are i or j.

    Fix: Select only the examples belonging to the chosen pair for an All-Pairs task.

Check Your Understanding

MEDIUM

A problem has three possible classes. Explain how One-Versus-All and All-Pairs would organize its binary training tasks. Then explain what kind of results would be considered together when making a multiclass prediction.

Hints
  • For One-Versus-All, identify how many class-versus-rest tasks are created.
  • For All-Pairs, list the kinds of class pairs that can be selected.
  • Keep the training structure separate from the prediction process.

What do you think happens?

For a three-class problem, which approach creates a binary task whose training examples contain only classes 1 and 3?

  • One-Versus-All
  • All-Pairs
  • Both approaches in the same form
Reveal answer

Answer: All-Pairs

The All-Pairs Approach creates a binary training set for every pair of classes, and the task for classes 1 and 3 contains only examples whose original labels are 1 or 3.

Key Takeaways

  1. X is the domain set of objects being studied; Y is the label set of possible labels.
  2. A domain point is usually represented by a feature vector, and features are not labels.
  3. In the source papaya example, X is the set of all papayas and Y is {0, 1}.
  4. A multiclass problem has more than two target classes and can be reduced to several binary problems.
  5. One-Versus-All trains one class-versus-rest classifier per class, while All-Pairs trains one classifier for every pair of classes.
  6. The two approaches differ both in which examples each binary task uses and in how the binary results are considered together for multiclass prediction.

Key Takeaways

  • The domain set X contains the objects under study, while the label set Y contains possible labels.
  • Feature vectors represent domain points; their features should not be confused with labels.
  • Multiclass prediction can be reduced to several binary classification tasks.
  • One-Versus-All creates one class-versus-rest task for each class.
  • All-Pairs creates one task for every pair of classes and uses those pairwise results together for prediction.