Concepts / Statistical Learning Framework

Statistical Learning Framework

X names the collection of objects being studied, while Y names the collection of possible labels.

  • Programming

Starting with Two Questions

A statistical learning problem begins by separating two questions: What objects are being studied, and what labels could describe those objects? The collection of objects is called the domain set and is written as X. The collection of possible labels is called the label set and is written as Y. This separation is the foundation of the framework.

containscontainscontainsXdomain setpapayasobjects being studied0possible labelYlabel set1possible label
How is an object in X different from a possible label in Y, and where does each appear in the learning setup?

Tracing a Domain Point

Take the papaya example. X is the set of all papayas, so one particular papaya is an individual instance, or domain point, in X. The domain point is the object being studied. It is not itself one of the labels.

considered withxone papaya in Xya possible label in Y
How does a domain object x from X become considered together with a possible label y from Y?

The framework keeps two roles distinct: x identifies a domain object, while y identifies a possible label for that object. The object comes from X; the possible label comes from Y.

Representing Objects with Features

A domain point is usually represented by a feature vector. A feature vector is a vector whose positions contain feature values describing the domain point for the learning problem. In the papaya example, color and softness are features. They provide the representation of the papaya; they are not labels.

represented bycontains featurecontains featureone papayadomain pointfeature vectorrepresentation of thepapayafeature positioncolorfeature positionsoftness
How does one object in the domain become a vector of feature values, and what does each position represent?

Connecting X and Y

The domain set X contains the objects being studied. The label set Y contains the possible labels that may be used with those objects. A domain point x is taken from X, and a possible label y is taken from Y.

containscontainssupplies objectsupplies possible labelXobjects being studiedxdomain pointlearning setupobject considered withlabelYpossible labelsypossible label
What does each set contain, and how are X and Y connected in the statistical learning framework?

Classifying the Parts of the Papaya Example

Identify what belongs to X, what belongs to Y, and what belongs to the feature representation.

Domain set: X is the set of all papayas. Every papaya, including one particular papaya, is on the domain side.

Feature representation: Color and softness are features of a domain point. They belong in the feature representation of the papaya.

Label set: Y is {0, 1}. The values 0 and 1 are the possible labels in this binary classification example.

Pairing: A particular papaya from X is considered together with a possible label from Y. The papaya remains the object, while the selected value remains the label.

Papayas are objects in X, color and softness are features describing a domain point, and 0 and 1 are labels in Y.

Mistakes in Classification

  • Treating a papaya as a label

    A papaya is an object being studied, so it belongs to X.

    Fix: Place all papayas, including the individual papaya, in the domain set X.

  • Treating color and softness as the two labels

    Color and softness are features used to represent a domain point.

    Fix: Keep color and softness in the feature representation and place 0 and 1 in Y.

  • Mixing the domain set with the label set

    The framework uses X for studied objects and Y for possible labels.

    Fix: Ask separately what object is being studied and what labels could describe it.

ItemRolePapaya example
XDomain set containing objects being studiedAll papayas
A domain pointOne individual instance in XOne particular papaya
Feature vectorRepresentation of a domain pointFeatures including color and softness
YLabel set containing possible labels{0, 1}
A labelA possible value from Y0 or 1

Practice Check

What do you think happens?

In the papaya example, which category does each item belong to: a particular papaya, color, softness, 0, and 1?

  • The papaya is in X; color and softness are features; 0 and 1 are in Y.
  • The papaya is in Y; color and softness are labels; 0 and 1 are in X.
  • All five items are labels in Y.
Reveal answer

Answer: The papaya is in X; color and softness are features; 0 and 1 are in Y.

The domain side contains the objects being studied. Features represent a domain point, while the label side contains the possible labels.

EASY

Explain in your own words why a feature vector does not become a label merely because it is used in the learning problem.

Hints
  • Identify what the feature vector represents.
  • Identify where the possible labels come from.
  • Use color and softness from the papaya example.

Key Takeaways

  1. X is the domain set: the collection of objects being studied.
  2. An individual object in X is a domain point and is usually represented by a feature vector.
  3. Features such as color and softness describe a domain point; they are not labels.
  4. Y is the label set: the collection of possible labels.
  5. In the papaya binary classification example, Y is {0, 1}, and a domain object from X is considered together with a possible label from Y.

Key Takeaways

  • The domain set X contains the objects under study, while the label set Y contains possible labels.
  • A domain point is one individual object in X and is usually represented by a feature vector.
  • In the papaya example, color and softness are features of the object, not labels.
  • For binary classification, the example label set is Y = {0, 1}.
  • A learning setup considers an object from X together with a possible label from Y.