Multiclass Classification and Ranking
Hinge loss is a classification loss defined in the source for binary classification.
Beyond Two Labels
Many learning problems begin with labeled examples. A learner uses those examples to produce a function that can make predictions for new inputs. In binary classification, the possible outputs are limited to two classes. Multiclass classification extends this task: the prediction must be one label from a larger finite set. Ranking is important because a multiclass predictor must account for several possible classes when deciding which label to return.
From Features to a Ranked Choice
A document topic task illustrates the multiclass setting. Each document can be represented as a feature vector, and the training sample contains pairs consisting of a feature vector and its correct label. The learner uses those examples to produce a function from the domain set to the label set. When a new document arrives, the function returns one topic from the larger finite collection.
The diagram represents the decision structure rather than a particular scoring formula. The important point is that several labels are possible, yet the learned function ultimately returns one label. The output set is therefore a larger finite collection of topics instead of a pair of alternatives.
Binary Hinge Loss
Hinge loss is a classification loss defined in the source for binary classification. Its binary expression is max { 0, 1 - y 〈w, x〉 }, where the outer maximum selects the larger of 0 and the quantity 1 - y 〈w, x〉.
The expression is specifically a binary definition. It describes a setting in which the prediction problem has two possible classes and uses y, w, and x in the displayed expression. It should not be presented as the complete loss formula for a multiclass predictor.
max { 0, 1 - y 〈w, x〉 }
Evaluating the binary expression
Suppose y 〈w, x〉 = 0.5. Evaluate max { 0, 1 - y 〈w, x〉 }.
Substitute: The expression becomes max { 0, 1 - 0.5 }.
Calculate the inner quantity: The quantity 1 - 0.5 equals 0.5.
Apply the maximum: The larger of 0 and 0.5 is 0.5.
The binary hinge-loss value for this generated numerical example is 0.5.
Why Multiclass Generalization Matters
A multiclass predictor faces more than one incorrect alternative. Because the predictor must account for multiple possible classes, binary hinge loss cannot simply be reused and called the full multiclass loss. Generalized hinge loss extends the hinge-loss idea to multiclass predictors, but the supplied source introduces that generalization without displaying its complete expression.
| Binary hinge loss | Generalized multiclass hinge loss |
|---|---|
| Defined in the source for binary classification | Extends hinge loss to multiclass predictors |
| Uses the displayed expression max { 0, 1 - y 〈w, x〉 } | Requires a complete multiclass expression before numerical calculation |
| Concerns a classification problem with two possible classes | Must account for multiple possible classes |
Classification and Regression
The output set distinguishes multiclass classification from regression. In multiclass classification, the learner chooses one label from a larger finite set, such as a document topic. In regression, the learner predicts a real-valued target by learning a functional relationship between inputs and outputs.
| Task | Output set | Source example |
|---|---|---|
| Multiclass classification | A larger finite set of labels | A document topic |
| Regression | The real numbers | A baby's weight measured in grams |
The target set, rather than the input format alone, identifies the type of learning task.
For the birth-weight example, each input contains three ultrasound measurements: head circumference, abdominal circumference, and femur length. The domain set is therefore a subset of R3. The target set is the real numbers, with weight measured in grams. In this setting, target is more appropriate than label because the output is a real number rather than a discrete topic.
Framing the Learning Problem
To frame a learning problem, identify five connected elements: the domain set of possible inputs, the representation of each input, the output set, the training pairs, and the learned function. The training data form a finite sequence of input-output pairs. The learner then produces a function from the domain set to the appropriate label set or target set.
- Identify the domain set containing the possible inputs.
- State how each input is represented, such as a feature vector.
- Identify the output set: a larger finite label set for multiclass classification or the real numbers for regression.
- Describe the finite sequence of training input-output pairs.
- State that the learner produces a function from the domain set to the selected output set.
Classifying document topics
Frame a document-topic learning problem using the learning framework.
Domain and representation: The inputs are documents represented as feature vectors.
Output set: The outputs belong to a larger finite set of document-topic labels.
Training examples: The learner receives a finite sequence of feature-vector and topic-label pairs.
Learned function: The learner produces a function from the domain set of documents to the topic-label set.
This is a multiclass classification problem because the output is one topic from a larger finite set.
Mistakes to Avoid
Calling max { 0, 1 - y 〈w, x〉 } the full multiclass hinge-loss formula.
The source defines that displayed expression for binary classification and introduces generalized hinge loss separately for multiclass predictors.
Fix:
Identify the expression as binary hinge loss and obtain the explicit generalized multiclass formula before calculating a multiclass value.Treating multiclass classification as if it still had only two possible labels.
Multiclass classification predicts one label from a larger finite set.
Fix:
State the full finite label set or describe it as a larger finite collection of possible labels.Calling a real-valued prediction a label.
Regression predicts a real-valued target, such as weight measured in grams.
Fix:
Use target for the real-valued output and label for a discrete classification output.Leaving out the output set when framing a learning problem.
The output space is the defining difference between the classification and regression examples.
Fix:
Always identify the domain, representation, output set, training pairs, and learned function.
Check Your Understanding
A learner receives feature-vector and label pairs for documents, where each label is one topic from a finite collection. Is this binary classification, multiclass classification, or regression? Then explain why the binary hinge-loss expression alone is insufficient for calculating a multiclass loss.
Hints
- Inspect the size and kind of the output set.
- Compare the document-topic task with the real-valued birth-weight task.
- Remember that the supplied binary expression is not the complete generalized multiclass formula.
What do you think happens?
A task predicts one topic from a larger finite set of document topics. What kind of learning task is it?
Reveal answer
Answer: Multiclass classification
The output is one label from a larger finite set. Regression instead predicts a real-valued target.
Key Takeaways
- Hinge loss is defined in the source for binary classification as max { 0, 1 - y 〈w, x〉 }.
- Generalized hinge loss extends the hinge-loss idea to multiclass predictors that must account for multiple possible classes.
- The binary hinge-loss expression is not the full multiclass formula; an explicit generalized formula is required before calculating a numeric multiclass loss.
- Multiclass classification returns one label from a larger finite set, while regression predicts a real-valued target.
- A learning problem is framed by identifying its domain, input representation, output set, training pairs, and learned function.