Expected Loss
Generalized loss functions provide a broad way to assign nonnegative values to model-domain pairs.
From Behavior to Cost
A learning model is useful only when we can describe how costly its behavior is on a domain example. Expected loss provides a general framework for doing that. Instead of requiring one particular learning task or one particular error measure, it lets us define a rule that accepts a model and a domain element, then returns a nonnegative real number.
The central distinction is simple: loss evaluates one model-domain pair, while risk summarizes expected loss across domain elements distributed according to D.
The Two Inputs
A generalized loss function is written as ℓ : H × Z → R+. This notation identifies both the inputs and the output. The first input is a hypothesis or classifier h from H. The second input is a domain element z from Z. After receiving that pair, the function returns a nonnegative real number: the loss assigned to that classifier on that domain element.
One Loss Evaluation
Evaluating One Pair
Suppose h is a classifier in H and z is a domain element in Z. A generalized loss function assigns the pair (h, z) the value 4.
Identify the first input: h is the hypothesis or classifier supplied to the loss function.
Identify the second input: z is the domain element supplied alongside h.
Read the output: The value 4 is the loss for this one specific classifier-domain pair. It is nonnegative, as required by the mapping into R+.
The evaluation describes ℓ(h, z), one loss value. It does not yet describe the classifier's risk.
A single loss value answers a narrow question: how costly is this particular classifier on this particular domain element? It does not summarize performance over a distribution of domain elements.
From Loss to Risk
Risk is the broader assessment of a classifier. It is the expected loss of a classifier h in H with respect to a probability distribution D over Z. In other words, risk combines the losses that h receives on domain elements while taking into account how those domain elements are distributed according to D.
| Quantity | What it evaluates | What it depends on |
|---|---|---|
| Loss | One classifier-domain pair | h, z, and the loss function |
| Risk | Expected performance of one classifier | The classifier, the loss function, and distribution D over Z |
Changing the Distribution
Same Classifier, Different Risk
A classifier h has loss 2 on z1 and loss 8 on z2. Consider two possible distributions over the same domain: D1 places more probability on z1, while D2 places more probability on z2.
Keep the classifier fixed: The classifier h has not changed, and the loss function has not changed.
Keep the individual losses fixed: The loss on z1 remains 2, and the loss on z2 remains 8.
Change the distribution: D1 and D2 assign different probability distributions over the domain elements.
Compare the expected losses: The distribution emphasizing z1 produces a lower expected loss than the distribution emphasizing z2, because h has the lower loss on z1.
The risk can change when D changes, even though the classifier and the loss function remain unchanged.
Common Interpretation Errors
Treating a loss value as the classifier's risk.
That value describes only one classifier-domain pair. Risk is expected loss with respect to a distribution D over Z.
Fix:
Ask whether the quantity concerns one z or aggregates losses under D.Forgetting that the loss function takes two inputs.
The mapping is from H × Z, so it accepts both a hypothesis or classifier and a domain element.
Fix:
Identify h and z before interpreting the output.Assuming that changing D cannot affect risk.
Risk is an expectation with respect to D, so changing the distribution used for that expectation can change the result.
Fix:
Treat D as part of the context in which risk is measured.Allowing the generalized loss output to be negative.
The mapping ends in R+, the nonnegative real numbers.
Fix:
Check that every assigned loss is nonnegative.
Check Your Understanding
A loss function is written as ℓ : H × Z → R+. For a chosen classifier h, suppose the domain contains z1 and z2, and the loss on z1 is smaller than the loss on z2. Explain which distribution is likely to give h the lower risk: one that emphasizes z1 or one that emphasizes z2. Then explain why the answer would change if the losses were reversed.
Hints
- Risk is expected loss with respect to D.
- Compare which domain element receives more probability.
- Keep the classifier and loss function fixed while reasoning about the change in D.
What do you think happens?
If the classifier and loss function stay fixed but the probability distribution D over Z changes, can the classifier's risk change?
Reveal answer
Answer: Yes
Risk is expected loss with respect to D. Changing the distribution used for the expectation can change the risk even when the classifier and loss function remain the same.
Key Takeaways
- A generalized loss function maps H × Z to the nonnegative real numbers.
- Its two inputs are a hypothesis or classifier h and a domain element z.
- Evaluating the loss function once gives the loss for one model-domain pair.
- Risk is the expected loss of a classifier with respect to a probability distribution D over Z.
- Changing D can change risk even when the classifier and loss function stay fixed.
Key Takeaways
- Generalized loss functions provide a broad way to assign nonnegative values to model-domain pairs.
- The notation ℓ : H × Z → R+ shows that the inputs are a hypothesis from H and a domain element from Z.
- A loss value evaluates one specific pair, whereas risk summarizes expected loss for a classifier.
- Risk depends on the probability distribution D over Z, so changing D can change the risk.