Concepts / Classifiers

Classifiers

Generalized loss functions provide a broad way to assign nonnegative values to model-domain pairs.

  • Programming

From Predictions to Cost

A classifier produces behavior on domain examples, but describing that behavior is not enough for learning theory. We also need a way to say how costly a particular classifier is on a particular domain example. A generalized loss function supplies that description. It accepts a hypothesis or model and a domain element, then returns a nonnegative real number.

A generalized loss function is written as ℓ : H × Z → R+. Its inputs are a hypothesis h from H and a domain element z from Z. Its output is a nonnegative real number representing the loss assigned to that pair.

model inputdomain inputHypothesis hh ∈ HLoss ℓ(h,z)a nonnegative real numberDomain element zz ∈ Z
What enters a generalized loss function, and what does it produce?

One Loss Versus Overall Risk

A single evaluation of ℓ describes one model-domain pair. Risk asks a broader question: how much loss should we expect from a chosen classifier when domain elements are distributed according to a probability distribution D over Z? Thus, risk is the expected loss of the classifier with respect to D.

Risk of h under D = expected loss of h when z is distributed according to D
evaluate h on z1evaluate h on z2weighted by DDomain element z1probability under DLoss valuesℓ(h,z) across ZRisk of hexpected loss under DDomain element z2probability under D
How do losses on domain elements combine under D to produce a classifier's risk?
evaluatetake expectationOne pair (h,z)one loss valueLoss ℓ(h,z)nonnegative real numberDistribution D overZmany possible domainelementsRisk of hexpected loss under D
What is the difference between evaluating one pair and averaging expected loss over domain elements?

Why the Distribution Matters

Suppose a fixed classifier and a fixed loss function are evaluated on domain elements z1 and z2. What can change if the distribution D changes?

Keep the classifier fixed: The selected classifier remains the same.

Keep the loss rule fixed: The generalized loss function still assigns a nonnegative value to each model-domain pair.

Change D: The probabilities used in the expectation change.

Recompute expected loss: Because the expectation is now taken with respect to a different distribution, the resulting risk can change.

Risk can change when D changes even though the classifier and loss function remain unchanged.

Three Regions on the Real Line

A 3-piece classifier operates on the real line. It is specified by two real thresholds, theta_1 and theta_2, satisfying theta_1 < theta_2, together with a binary sign parameter b in {+1, -1}. The ordered thresholds divide the real line into three intervals. The sign parameter determines the classifier's binary prediction pattern across those regions.

region predictionregion predictionregion predictionsets binary signLeft intervalx < theta_1Sign b+1 or -1Middle intervaltheta_1 to theta_2Binary predictionspattern across threeregionsRight intervalx > theta_2
How do two ordered thresholds divide the real line, and where does the sign parameter enter?

Reading a 3-Piece Classifier's Structure

Consider parameters theta_1 = 2, theta_2 = 5, and b = +1. What structural information does this specify?

Check the threshold order: The thresholds satisfy 2 < 5, so they are valid ordered thresholds.

Locate the boundaries: The real line is divided at 2 and 5.

Identify the three regions: The classifier has a region to the left of the first threshold, a region between the thresholds, and a region to the right of the second threshold.

Apply the sign parameter: The value b = +1 selects the binary sign orientation used for the predictions in those regions.

The parameters specify two ordered boundaries, three real-line regions, and a binary prediction orientation.

Stumps and Weak Learning

A decision stump is simpler than a 3-piece classifier. It uses one threshold theta and a sign parameter b, rather than two ordered thresholds. Consequently, a stump divides the real line using one boundary, while a 3-piece classifier divides it using two boundaries and has three prediction regions.

divides real linedivide real lineOne threshold thetaone boundaryTwo regionsbinary predictionsTwo thresholdstheta_1, theta_2theta_1 < theta_2Three regionsbinary predictions
What changes when the prediction structure moves from one threshold to two?

In this example, decision stumps form the learner class B, while 3-piece classifiers form the larger target class H. The point of calling stumps a weak learner is not that a stump has the same structure as a 3-piece classifier. It does not. Rather, the simpler stump family can still provide a learning guarantee about the more structured target family.

simpler class used for learningchoose empirical risk minimizerachieves3-piece class Htwo thresholdsStump class Bone thresholdERM_Bbest stump in BWeak-learningguaranteegamma = 1/12
How can a simpler one-threshold learner support learning about a two-threshold classifier family?

The statement that ERM_B is a gamma-weak learner for H with gamma equal to 1/12 means that empirical risk minimization over the stump class B provides a guarantee whose gap from the best classifier in the target class H is at most 1/12 in risk.

Common Misreadings

  • Treating the loss value and risk as the same quantity.

    ℓ(h,z) evaluates one hypothesis-domain pair, whereas risk is expected loss under D.

    Fix: Use loss for one pair and risk for the distribution-based expectation.

  • Forgetting that D is part of the risk question.

    Changing D can change the expectation.

    Fix: Always specify the distribution with respect to which risk is measured.

  • Describing a 3-piece classifier with only one threshold.

    The 3-piece class is specified by two ordered thresholds and a binary sign parameter.

    Fix: Check for theta_1, theta_2 with theta_1 < theta_2, plus b in {+1, -1}.

  • Assuming a decision stump and a 3-piece classifier have the same structure.

    A stump uses one threshold, while the 3-piece classifier uses two.

    Fix: Compare the number of thresholds and resulting regions.

  • Interpreting weak learner as perfect learner.

    The guarantee allows a risk gap of gamma, which here is 1/12.

    Fix: Interpret the guarantee as being within 1/12 in risk of the best classifier in H.

Check Your Understanding

MEDIUM

A learning setup has a hypothesis class H, a domain Z, and a probability distribution D over Z. Explain, in your own words, the difference between evaluating ℓ(h,z) and evaluating the risk of h under D. Then compare the structures of a decision stump and a 3-piece classifier.

Hints
  • Start by identifying the two inputs to ℓ.
  • Ask whether the calculation concerns one domain element or an expectation over domain elements.
  • Count the thresholds in each classifier family.

What do you think happens?

If the classifier and generalized loss function stay fixed but D changes, should the risk necessarily stay fixed?

  • Yes, because the classifier did not change
  • No, because risk is an expectation with respect to D
Reveal answer

Answer: No, because risk is an expectation with respect to D.

Changing the distribution changes the weighting used in the expected loss, so it can change risk even when the classifier and loss function are unchanged.

Key Takeaways

  1. A generalized loss function maps a hypothesis-domain pair from H × Z to a nonnegative real number.
  2. Loss evaluates one pair, while risk is the expected loss of a classifier under a distribution D over Z.
  3. A 3-piece classifier uses two ordered real thresholds and a binary sign parameter to divide the real line into three regions.
  4. A decision stump uses one threshold and is therefore structurally simpler than a 3-piece classifier.
  5. ERM_B being a 1/12-weak learner for H means that learning over the stump class B has a risk guarantee within 1/12 of the best classifier in H.

Key Takeaways

  • Generalized loss functions assign nonnegative values to pairs consisting of a hypothesis and a domain element.
  • Risk is expected loss under a distribution D, not one individual loss value.
  • Two ordered thresholds create the three regions of a 3-piece classifier.
  • Decision stumps use one threshold and can still act as weak learners for the richer 3-piece class.
  • The value gamma = 1/12 describes the allowed risk gap in the weak-learning guarantee.