Concepts / 0-1 Loss

0-1 Loss

The multiclass categorization goal is to learn h : X → [k].

  • Programming

The Prediction Goal

A multiclass categorization problem asks us to learn a predictor h : X → [k]. The predictor takes an input from X and returns one output from [k]. This describes the prediction task: construct a rule that turns inputs into class predictions.

The predictor answers the question: given an input x from X, which one label in [k] should be returned?

From Inputs to Labels

The two sides of h : X → [k] describe different parts of the prediction task. X is the input side: it supplies the objects that the predictor receives. [k] is the output side: it supplies the possible class labels that the predictor may return. The predictor's role is to connect one input with one output label.

hhx₁input from X1label in [k]x₂input from X3label in [k]
How does a predictor h map each input x in X to one label in the output set [k]?

Tracing One Prediction

Consider an input x₁ from X and a predictor h whose output set is [k]. What does the predictor do?

Receive: The predictor receives x₁, an input from X.

Apply h: The predictor applies the mapping h to that input.

Return: The predictor returns one class label from [k], such as label 1 in this generated illustration.

The prediction task is complete when h returns a label. At this point, we know what was predicted, but not whether the prediction is good.

Prediction and Evaluation

Producing a label and evaluating that label are separate tasks. The predictor h produces an output in [k]. An evaluation rule then considers whether that output is correct for the example. The source description emphasizes this separation: the mapping tells us what the predictor does, while an evaluation rule tells us how to discuss the quality of its output.

apply hcompareevaluatexinput from Xh(x)label from [k]prediction and truelabelevaluation step0-1 lossperformance outcome
How is producing a label h(x) different from comparing that label with the true label to measure performance?

Correct and Incorrect Outcomes

The 0-1 loss evaluates a multiclass prediction through a correct-or-incorrect outcome. When the predicted class agrees with the true class, the loss records the correct outcome as 0. When the predicted class does not agree with the true class, the loss records the incorrect outcome as 1. The important idea is that the loss evaluates the returned label; it does not produce the label.

with true label 2matcheswith true label 3does not matchpredicted label 2before evaluation0correct outcometrue label 2matching case1incorrect outcometrue label 3nonmatching case
What changes in the evaluation when the predicted label matches the true label versus when it does not?

Two Evaluation Cases

A predictor returns label 2 for an input. Compare two possible true labels.

Matching label: If the true label is also 2, the prediction agrees with the true label, so the 0-1 loss records 0.

Different label: If the true label is 3, the prediction does not agree with the true label, so the 0-1 loss records 1.

The predictor's returned label is the same kind of object in both cases. The evaluation outcome changes because the relationship between the prediction and the true label changes.

Aggregating Performance

A single prediction produces one evaluation outcome. Across a collection of examples, these individual correct-or-incorrect outcomes provide a way to discuss how well the multiclass predictor performs. This is the performance perspective needed after the predictor has produced labels.

apply hevaluateconsider togethercollection ofexamplesinputs and true labelsclass predictionsoutputs from h0 or 1 outcomescorrect or incorrectpredictor performanceevaluation across examples
How do individual correct-or-incorrect outcomes combine to measure how well a multiclass predictor performs?

The 0-1 loss shifts attention from an individual returned label to the predictor's pattern of correct and incorrect outcomes across examples.

Connection to PAC Learnability

PAC learnability for multiclass predictors is considered with respect to the 0-1 loss. This means that the learning discussion evaluates a predictor through the correct-or-incorrect outcomes produced on examples. The predictor h remains the rule that maps inputs in X to labels in [k]; the 0-1 loss is the rule used to assess how well that predictor performs.

Common Mistakes

  • Treating X as the set of possible class labels.

    X is the input side of the mapping, while [k] is the output side.

    Fix: Read h : X → [k] from left to right: an input from X is mapped to an output in [k].

  • Treating the prediction h(x) as the evaluation result.

    h(x) is the returned class label. The 0-1 loss evaluates that label.

    Fix: First identify the predicted label, then use the evaluation rule to determine the correct-or-incorrect outcome.

  • Describing multiclass categorization only as an evaluation task.

    The central task is to learn a predictor that turns inputs from X into class predictions in [k].

    Fix: State the mapping goal before discussing its performance.

Check Your Understanding

EASY

A predictor h receives an input x from X and returns a label in [k]. Explain, in two separate statements, what the predictor does and what the 0-1 loss evaluates.

Hints
  • Use the roles of X and [k] in h : X → [k].
  • Keep the returned label separate from the correct-or-incorrect evaluation outcome.

What do you think happens?

A predictor returns label 1. Has its performance already been determined?

  • Yes, because h has returned a label.
  • No, because the returned label must be evaluated against the true label.
Reveal answer

Answer: No, because the returned label must be evaluated against the true label.

Producing a class prediction and evaluating that prediction are separate tasks. The 0-1 loss supplies the evaluation perspective.

Key Takeaways

  1. Multiclass categorization aims to learn a predictor h : X → [k].
  2. X supplies inputs to the predictor, and [k] supplies the possible output labels.
  3. The predictor produces a label; the 0-1 loss evaluates whether that label is correct.
  4. Across examples, correct-or-incorrect outcomes describe predictor performance.
  5. PAC learnability for multiclass predictors is considered with respect to the 0-1 loss.

Key Takeaways

  • The goal is to learn a mapping h from inputs in X to labels in [k].
  • Prediction and evaluation are different stages.
  • The 0-1 loss records whether a predicted label agrees with the true label.
  • PAC learnability uses this loss to discuss the performance of multiclass predictors.