Concepts / Learning Problems

Learning Problems

X names the collection of objects being studied, while Y names the collection of possible labels.

  • Programming

Start with Two Collections

A learning problem separates two questions: what objects are being studied, and what labels could describe those objects? The collection of studied objects is the domain set X. The collection of possible labels is the label set Y. Keeping X and Y separate is the foundation for reading the rest of the notation correctly.

containscontainscontainsDomain set Xobjects being studiedpapayaone domain point0possible labelLabel set Ypossible labels1possible label
What is the difference between an object in the domain set X and a possible label in the label set Y?

From Object to Feature Vector

An individual instance in X is called a domain point. A domain point is usually represented by a vector of features. In the papaya example, color and softness are features of the domain point. They describe the object used in the learning problem; they are not labels and do not replace the domain object with a label.

represented bycontainscontainspapayadomain pointfeature vectorrepresentationcolorfeature positionsoftnessfeature position
How does one domain object become a feature vector, and what does each position in that vector represent?

Pairing an Object with a Label

The learning problem considers a domain object together with a possible label. Begin with one papaya from X. Represent that papaya using its features, such as color and softness. Then interpret the possible outcome using a value from Y. For the papaya example, Y is {0, 1}, so 0 and 1 are the allowed labels. The papaya remains the domain object, while the selected value comes from the label set.

representconsider withpapayaobject from Xcolor and softnessfeature representation0 or 1possible label from Y
How does a domain object x from X become considered together with a possible label y from Y?

Classifying the Papaya Information

Place each item on the correct side of the learning problem: X, Y, or the feature representation of a domain point.

Identify X: X is the set of all papayas. One particular papaya is an individual domain point in X.

Identify Y: Y is {0, 1}. The values 0 and 1 are possible labels.

Identify features: Color and softness are features used to represent a papaya.

Keep the roles separate: The papaya is the object being studied, color and softness describe that object, and 0 or 1 supplies a possible label.

The domain side contains papayas, the label side contains 0 and 1, and color and softness belong to the feature representation.

Reading the Learning-Problem Tuple

Both named definitions use the same learning-problem notation: (H, Z, ℓ). The tuple is the learning problem that appears before the parameter pair is identified. In the supplied notation guide, H is the hypothesis set, Z is the example space, and ℓ is the loss. The tuple stays the same in both definitions; the distinction comes from the parameter names.

part ofpart ofHhypothesis set,Zexample space,ℓloss
What does each component of the tuple (H, Z, ℓ) contain, and how are the hypothesis set, example space, and loss connected?

When you see either named learning-problem definition, first locate (H, Z, ℓ). That tuple identifies the learning problem being discussed. Then inspect the parameter pair that follows or is named with the definition.

ρ, B versus β, B

DefinitionLearning-problem tupleParameter pairRecognition clue
Convex-Lipschitz-Bounded(H, Z, ℓ)ρ, BFirst parameter is ρ
Convex-Smooth-Bounded(H, Z, ℓ)β, BFirst parameter is β

The Convex-Lipschitz-Bounded definition names ρ and B. The Convex-Smooth-Bounded definition names β and B. Both use the same learning-problem tuple, and both include B. The safest classification method is therefore to inspect the first parameter: ρ identifies the Convex-Lipschitz-Bounded definition, while β identifies the Convex-Smooth-Bounded definition.

first parameter changesConvex-Lipschitz-Bounded(H, Z, ℓ); ρ, BConvex-Smooth-Bounded(H, Z, ℓ); β, B
What changes between the two learning-problem definitions when the parameter pair is ρ, B versus β, B?

Recognition Practice

EASY

A definition uses the learning problem (H, Z, ℓ) together with the parameter pair β, B. Which named learning-problem definition does the parameter pair identify?

Hints
  • The tuple is the same in both definitions.
  • Inspect the first parameter rather than B.
MEDIUM

Sort these items into the correct roles in the papaya example: all papayas, one papaya, color, softness, 0, and 1. Use the roles domain set, domain point, feature, and possible label.

Hints
  • The domain set contains the objects being studied.
  • The label set in this example is {0, 1}.
  • Color and softness describe a domain point.
  • Treating a papaya as a label.

    A papaya is an object in the domain set X, while the source example places 0 and 1 in the label set Y.

    Fix: Keep the studied objects on the X side and the possible label values on the Y side.

  • Treating color and softness as the two labels.

    Color and softness are features used to represent a domain point; 0 and 1 are the labels in the source example.

    Fix: Separate the feature representation from the label set.

  • Choosing a definition from the shared parameter B.

    Both named definitions include B.

    Fix: Inspect the first parameter: ρ indicates Convex-Lipschitz-Bounded, and β indicates Convex-Smooth-Bounded.

  • Assuming the two definitions use different learning-problem tuples.

    The source states that a learning problem in both definitions is written as (H, Z, ℓ).

    Fix: Treat the tuple as unchanged and compare the parameter pair.

Final Checklist

  1. X is the domain set: its elements are the objects being studied.
  2. A domain point is an individual object in X and is usually represented by a feature vector.
  3. Y is the label set: in the papaya binary-classification example, Y is {0, 1}.
  4. The learning problem in both named definitions is written as (H, Z, ℓ).
  5. ρ, B identifies the Convex-Lipschitz-Bounded parameter pair, while β, B identifies the Convex-Smooth-Bounded parameter pair.

Key Takeaways

  • The domain set X contains the objects being studied, while the label set Y contains possible labels.
  • A domain point can be represented by a feature vector; color and softness are features in the papaya example.
  • For binary classification in the source example, the label set is {0, 1}.
  • Both named learning-problem definitions use the tuple (H, Z, ℓ).
  • The first parameter distinguishes the definitions: ρ for Convex-Lipschitz-Bounded and β for Convex-Smooth-Bounded.