Statistical Learning Framework
X names the collection of objects being studied, while Y names the collection of possible labels.
Starting with Two Questions
A statistical learning problem begins by separating two questions: What objects are being studied, and what labels could describe those objects? The collection of objects is called the domain set and is written as X. The collection of possible labels is called the label set and is written as Y. This separation is the foundation of the framework.
Tracing a Domain Point
Take the papaya example. X is the set of all papayas, so one particular papaya is an individual instance, or domain point, in X. The domain point is the object being studied. It is not itself one of the labels.
The framework keeps two roles distinct: x identifies a domain object, while y identifies a possible label for that object. The object comes from X; the possible label comes from Y.
Representing Objects with Features
A domain point is usually represented by a feature vector. A feature vector is a vector whose positions contain feature values describing the domain point for the learning problem. In the papaya example, color and softness are features. They provide the representation of the papaya; they are not labels.
Connecting X and Y
The domain set X contains the objects being studied. The label set Y contains the possible labels that may be used with those objects. A domain point x is taken from X, and a possible label y is taken from Y.
Classifying the Parts of the Papaya Example
Identify what belongs to X, what belongs to Y, and what belongs to the feature representation.
Domain set: X is the set of all papayas. Every papaya, including one particular papaya, is on the domain side.
Feature representation: Color and softness are features of a domain point. They belong in the feature representation of the papaya.
Label set: Y is {0, 1}. The values 0 and 1 are the possible labels in this binary classification example.
Pairing: A particular papaya from X is considered together with a possible label from Y. The papaya remains the object, while the selected value remains the label.
Papayas are objects in X, color and softness are features describing a domain point, and 0 and 1 are labels in Y.
Mistakes in Classification
Treating a papaya as a label
A papaya is an object being studied, so it belongs to X.
Fix:
Place all papayas, including the individual papaya, in the domain set X.Treating color and softness as the two labels
Color and softness are features used to represent a domain point.
Fix:
Keep color and softness in the feature representation and place 0 and 1 in Y.Mixing the domain set with the label set
The framework uses X for studied objects and Y for possible labels.
Fix:
Ask separately what object is being studied and what labels could describe it.
| Item | Role | Papaya example |
|---|---|---|
| X | Domain set containing objects being studied | All papayas |
| A domain point | One individual instance in X | One particular papaya |
| Feature vector | Representation of a domain point | Features including color and softness |
| Y | Label set containing possible labels | {0, 1} |
| A label | A possible value from Y | 0 or 1 |
Practice Check
What do you think happens?
In the papaya example, which category does each item belong to: a particular papaya, color, softness, 0, and 1?
Reveal answer
Answer: The papaya is in X; color and softness are features; 0 and 1 are in Y.
The domain side contains the objects being studied. Features represent a domain point, while the label side contains the possible labels.
Explain in your own words why a feature vector does not become a label merely because it is used in the learning problem.
Hints
- Identify what the feature vector represents.
- Identify where the possible labels come from.
- Use color and softness from the papaya example.
Key Takeaways
- X is the domain set: the collection of objects being studied.
- An individual object in X is a domain point and is usually represented by a feature vector.
- Features such as color and softness describe a domain point; they are not labels.
- Y is the label set: the collection of possible labels.
- In the papaya binary classification example, Y is {0, 1}, and a domain object from X is considered together with a possible label from Y.
Key Takeaways
- The domain set X contains the objects under study, while the label set Y contains possible labels.
- A domain point is one individual object in X and is usually represented by a feature vector.
- In the papaya example, color and softness are features of the object, not labels.
- For binary classification, the example label set is Y = {0, 1}.
- A learning setup considers an object from X together with a possible label from Y.