Classification Loss Functions
Logistic loss evaluates how bad a prediction is for its true label.
Why a Prediction Needs a Loss
A logistic regression model produces h_w(x), a value in the interval [0, 1]. The model uses this value in relation to the two possible labels, y ∈ {−1, 1}. A prediction is not equally useful in every situation: it may support the actual label or work against it. Logistic loss turns that relationship into a numerical measure of how bad the prediction is for its true label.
Logistic loss evaluates one prediction on one labeled example. Empirical Risk Minimization then combines those individual evaluations across the training set.
Reading the Signed Score
The logistic loss is defined by log(1 + exp(−y〈w, x〉)). The important quantity inside this expression is the signed product y〈w, x〉. It evaluates the model's score relative to the actual label, rather than inspecting 〈w, x〉 by itself.
The sign of y〈w, x〉 tells us whether the model's score agrees with the label. A positive value indicates that the score supports the true label. A negative value indicates disagreement with the label and leads to higher loss. As the signed quantity becomes less favorable and eventually negative, the logistic loss increases.
The diagram shows direction rather than particular numerical loss values. Its central lesson is that the loss evaluates the score relative to the label. The same raw score 〈w, x〉 can therefore be interpreted differently when the actual label changes, because the product y〈w, x〉 changes with y.
From Support to Contradiction
Comparing two labeled examples
Consider two examples evaluated by the same logistic-regression model. In the first, the signed quantity y〈w, x〉 is strongly positive. In the second, it is strongly negative. Which example receives the larger logistic loss?
First example: A strongly positive y〈w, x〉 means that the model's score supports the true label. The corresponding logistic loss is lower than it would be for a less favorable signed quantity.
Second example: A strongly negative y〈w, x〉 means that the score disagrees with the true label. Negative values lead to higher logistic loss.
Comparison: The second example receives the larger loss because its signed quantity is less favorable and negative.
Logistic loss is lower when the signed score supports the true label and higher when the signed score contradicts it.
Averaging Loss with ERM
The loss for one example gives a local view of model quality. Empirical Risk Minimization, or ERM, provides the global view: it minimizes the average logistic loss over all training examples. The training process therefore evaluates each labeled example, combines those loss values through their average, and seeks model parameters that minimize that average.
This pipeline connects two levels of reasoning. First, logistic loss asks how bad one prediction is for its true label. Then ERM aggregates those individual judgments so that the model is evaluated over the entire training set rather than on only one example.
Common Interpretation Mistakes
Treating a positive raw score 〈w, x〉 as automatically favorable.
Logistic loss evaluates the score relative to the actual label through y〈w, x〉.
Fix:
Determine the sign of y〈w, x〉 before interpreting the prediction.Assuming that disagreement with the label should produce a smaller loss.
Negative values indicate disagreement with the label and lead to higher loss.
Fix:
Associate negative signed quantities with stronger penalty.Confusing one example's loss with the ERM objective.
ERM minimizes the average logistic loss over all training examples.
Fix:
Separate the local loss for one example from the average loss across the training set.
Check Your Reasoning
Two predictions have signed quantities y〈w, x〉 with opposite signs. Prediction A has a positive signed quantity, while Prediction B has a negative signed quantity. Without calculating numerical values, identify which prediction receives the higher logistic loss and explain why.
Hints
- Start by interpreting the sign of each signed quantity.
- Recall how logistic loss responds when the score disagrees with the true label.
Explain the difference between these two statements: logistic loss evaluates one labeled example, and ERM minimizes average logistic loss.
Hints
- Identify the local object being evaluated in the first statement.
- Identify what is aggregated in the second statement.
Key Takeaways
- Logistic loss measures how bad a logistic-regression prediction is for its true label.
- Its defining expression is log(1 + exp(−y〈w, x〉)).
- The sign of y〈w, x〉 matters: negative values indicate disagreement with the label and lead to higher loss.
- Predictions that support the true label receive lower loss than predictions that contradict it, in the directional sense described by the source.
- Empirical Risk Minimization minimizes the average logistic loss over the training set.
Key Takeaways
- Logistic loss evaluates the quality of a prediction relative to its true label.
- The signed product y〈w, x〉 is the key quantity because it incorporates both the model score and the actual label.
- A less favorable or negative signed quantity produces higher logistic loss.
- ERM uses the average of the individual logistic losses to evaluate and optimize performance across the training set.