Concepts / Multiclass Linear Predictors

Multiclass Linear Predictors

Multiclass prediction can be reduced to binary classification, but the reduction must include a rule for combining binary outputs.

  • Programming

From Two Labels to Many

Binary classification chooses between two possible labels. Multiclass categorization chooses one class from several possible target classes. Formally, a multiclass predictor maps an instance space X to a finite set of categories Y. For example, the instances might be documents and the categories might be their possible topics, or the instances might be images and the categories might be the possible objects appearing in them.

classifycategorizeInstanceTwo labelsone selectedInstanceSeveral classesone selected
What changes when a prediction has two possible labels instead of several?

A multiclass problem is not solved merely by collecting several binary predictions. The system also needs a rule that combines those binary outputs into one multiclass decision.

Class-Specific Scores

A multiclass linear predictor can be understood as comparing class-specific scores for one input. Each possible class contributes a score, and the prediction is the class associated with the largest score. This gives the learner a direct mental model for multiclass prediction: one input is evaluated against several class possibilities, then the scores are compared.

evaluateevaluateevaluateInputClass AscoreClass BscoreClass Clargest score
How do class-specific scores for one input determine the predicted class?

Selecting the Largest Class Score

An input is evaluated for three possible classes. The resulting scores are Class A: 0.42, Class B: 0.81, and Class C: 0.37. Which class is selected?

Compare: Place the three class-specific scores side by side.

Find the largest: The score for Class B is larger than the scores for Class A and Class C.

Select: Use the class associated with the largest score as the prediction.

The predicted class is Class B.

One-versus-All Reduction

One-versus-All, also called One-versus-Rest, contrasts each class with all remaining classes. For every possible class, a binary learner is trained to distinguish that class from the collection of other classes. A multiclass input is then passed through the class-specific binary learners, and their outputs are combined using a decision rule. The class with the strongest resulting score is selected.

evaluateevaluateevaluatescorescorescoreselect strongestInputClass A testA versus restCombine scoresPredicted classClass B testB versus restClass C testC versus rest
How does one multiclass input move through one binary classifier per class and become one final class?

One-versus-All for Three Topics

A document must be assigned to one of three topics: sports, science, or travel. Describe the One-versus-All binary tasks and use their scores to choose a topic.

Create binary tasks: Train one learner for sports versus all other topics, one for science versus all other topics, and one for travel versus all other topics.

Evaluate the document: The document is evaluated by all three binary learners.

Combine outputs: Compare the three resulting class scores using the multiclass decision rule.

Choose one class: Assign the document to the topic whose learner produces the strongest resulting score.

One-versus-All turns the three-class task into three binary contrasts and then needs a score-combination rule to produce one topic.

All-Pairs Voting

All-Pairs contrasts every pair of classes. Instead of comparing one class with all remaining classes, it creates a binary learner for each pair, such as Class A versus Class B and Class A versus Class C. For a new input, the pairwise learners produce binary outcomes. The system combines the resulting wins, and the class with the strongest overall vote is selected.

comparecomparecomparewinnerwinnerwinneraggregateInputA versus BClass votesPredicted classA versus CB versus C
How do pairwise comparisons combine their binary outcomes into one multiclass prediction?

Combining Pairwise Wins

Three classes, A, B, and C, are compared in every pair. An input is assigned the winner A in A versus B, the winner A in A versus C, and the winner C in B versus C. Which class receives the most wins?

List comparisons: The pairwise learners compare A with B, A with C, and B with C.

Count wins: Class A receives two wins. Class B receives no wins. Class C receives one win.

Aggregate: Use the accumulated wins as the multiclass decision signal.

Class A is selected because it receives the most pairwise wins.

When Reduction Breaks

A reduction method can perform poorly at the combination stage even when every binary learner has solved its own training problem correctly. The reason is that each binary learner is trained separately and does not directly know how its output will later participate in the multiclass decision. The binary tasks may therefore produce outputs that are individually reasonable but collectively difficult to combine.

evaluatemay producemay producecombinecombineInputBinary outputsConflicting votesOverall predictionpoor combinationAmbiguous scores
How can individually correct binary decisions still lead to conflicting votes, ambiguous scores, or an incorrect overall class?

Three Output Problems

Before comparing reduction methods, identify what the learner must produce. In multiclass categorization, an instance belongs to one of several possible target classes, so the predictor maps an instance to a finite set of categories. Structured output learning and ranking are different problem types rather than alternative names for this same multiclass task. The important diagnostic question is whether the desired result is one category from a finite set, an output with additional structure, or an ordering among possibilities.

selectproduceorderMulticlasscategorizationone categoryLearner outputStructured outputstructured resultRankingordered possibilities
How do multiclass categorization, structured output learning, and ranking differ in the kind of result the learner must produce?

State the output requirement before selecting One-versus-All or All-Pairs. A reduction is appropriate only after the task has been identified as multiclass categorization and the desired combination rule is clear.

Mistakes in Reduction Design

  • Treating several binary learners as a complete multiclass method.

    The reduction must include a rule for combining binary outputs.

    Fix: Define the score comparison or vote aggregation step explicitly.

  • Confusing One-versus-All with All-Pairs.

    One-versus-All contrasts each class with all remaining classes, whereas All-Pairs contrasts every pair of classes.

    Fix: Name the contrast used by each binary learner before discussing its outputs.

  • Assuming correct binary training guarantees correct multiclass prediction.

    The binary learners are trained without directly knowing how their outputs will later be combined.

    Fix: Evaluate the complete reduction, including the aggregation rule.

  • Choosing a reduction before identifying the output problem.

    Multiclass categorization, structured output learning, and ranking describe different output requirements.

    Fix: First state what the learner must produce.

Check Your Understanding

MEDIUM

A task has four possible target classes. You want to reuse binary learning algorithms. Describe how One-versus-All would create its binary tasks, then describe how All-Pairs would create its binary tasks. Finally, explain one reason the final multiclass predictor could perform poorly even if all binary learners were trained correctly.

Hints
  • For One-versus-All, describe one class versus all remaining classes.
  • For All-Pairs, describe comparisons between every pair of classes.
  • Include the missing combination or aggregation stage in your explanation.

What do you think happens?

A multiclass system has three class-specific scores for one input: Class A is 0.55, Class B is 0.31, and Class C is 0.72. Which class does the score-based decision rule select?

  • Class A
  • Class B
  • Class C
Reveal answer

Answer: Class C

Class C has the largest class-specific score.

Key Takeaways

  • Multiclass categorization maps an instance to one class from a finite set of possible categories.
  • One-versus-All contrasts each class with all remaining classes and combines the resulting class-specific outputs.
  • All-Pairs contrasts every pair of classes and uses the resulting wins to select a class.
  • A reduction method can fail during output combination even when its individual binary learners solve their own tasks correctly.
  • Before choosing a reduction, identify whether the learner must produce one category, a structured output, or an ordering.