Learning Models
Generalized loss functions provide a broad way to assign nonnegative values to model-domain pairs.
From Model Behavior to Cost
A learning model is useful only when we can describe how costly its behavior is on a domain example. A generalized loss function provides this description without requiring one particular learning task or one particular way to measure error. It takes a model and a domain element as inputs, then returns a nonnegative real number.
The central distinction in this topic is scope: a loss evaluates one model-domain pair, while risk evaluates a classifier across a distribution of domain elements.
The Model-Domain Pair
A generalized loss function is written as ℓ : H × Z → R+. Here, H and Z identify the input sets, and R+ identifies the output set of nonnegative real numbers.
The notation describes a mapping. The first input is a hypothesis, classifier, or model h from H. The second input is a domain element z from Z. Together, h and z form one model-domain pair. The function ℓ then assigns that pair one nonnegative loss value.
Tracing One Loss Evaluation
One Model-Domain Evaluation
Suppose a generated example uses a classifier hA, a domain element z1, and a generalized loss rule that assigns ℓ(hA, z1) = 2.
Choose the model: The first input is hA, a classifier in the model set H.
Choose the domain element: The second input is z1, a domain element in Z.
Apply the loss rule: The pair (hA, z1) is evaluated by ℓ.
Read the output: The assigned loss value is 2, which is a nonnegative real number.
This evaluation produces one individual loss value: ℓ(hA, z1) = 2. It does not yet describe the classifier's risk.
Risk Across a Distribution
Risk is the expected loss of a classifier h in H with respect to a probability distribution D over Z.
To find an individual loss, we supply one classifier and one domain element to ℓ. To discuss risk, we consider possible domain elements and the probability distribution D over Z. The risk combines the losses associated with those domain elements according to that distribution. Therefore, risk is a broader assessment than one evaluation of ℓ.
A Generated Risk Calculation
Consider a generated example with classifier hA and domain elements z1 and z2. Let D assign probability 0.8 to z1 and 0.2 to z2. Suppose ℓ(hA, z1) = 1 and ℓ(hA, z2) = 5.
Evaluate each domain point: The two individual loss values are 1 for z1 and 5 for z2.
Account for D: The loss at z1 is associated with probability 0.8, while the loss at z2 is associated with probability 0.2.
Combine the weighted losses: The expected loss is 0.8 × 1 + 0.2 × 5 = 1.8.
In this generated illustration, the risk of hA under D is 1.8. The individual losses are 1 and 5; the risk is the expected loss across the distribution.
Loss Versus Risk
| Quantity | What it considers | What it produces |
|---|---|---|
| Individual loss | One classifier-model and one domain element | One nonnegative loss value |
| Risk | A classifier and a probability distribution D over Z | Expected loss under D |
Mistakes in Scope
Treating ℓ(h, z) as the risk of h.
A single evaluation describes one loss value. Risk is the expected loss with respect to a probability distribution D over Z.
Fix:
Use loss for one model-domain pair and risk for the distribution-based expected loss.Forgetting one of the two inputs.
The generalized loss mapping takes both a model or classifier and a domain element.
Fix:
Check that the pair contains h from H and z from Z before applying ℓ.Assuming risk is fixed once the classifier is fixed.
Risk is an expectation with respect to D, so changing the distribution can change the expected loss.
Fix:
Always identify the distribution used when discussing a classifier's risk.
Check Your Understanding
A generalized loss function is applied to a classifier h and a domain element z. Explain what the resulting value represents. Then explain what additional idea is needed to move from that individual loss value to the classifier's risk.
Hints
- Start with the two inputs in H × Z.
- Risk is defined using an expected loss.
- Identify the probability distribution involved in that expectation.
Suppose the classifier and generalized loss function stay unchanged, but the probability distribution over Z changes. Can the risk change? Explain why.
Hints
- Risk is taken with respect to a distribution D over Z.
- Compare changing one individual loss evaluation with changing the distribution used for the expectation.
Key Takeaways
- The notation ℓ : H × Z → R+ describes a generalized loss function whose inputs are a model or classifier and a domain element.
- The output of the loss function is one nonnegative real value for one model-domain pair.
- Risk is the expected loss of a classifier with respect to a probability distribution D over Z.
- A change in D can change risk even when the classifier and loss function remain unchanged.
- Loss is local to one pair; risk is a broader, distribution-based assessment.
Key Takeaways
- A generalized loss function maps a pair from H × Z to a nonnegative real number.
- The two inputs are a classifier or model h and a domain element z.
- An individual loss evaluates one model-domain pair.
- Risk is the expected loss of a classifier under a probability distribution D over Z.
- Changing D can change risk without changing the classifier or the loss function.