Concepts / Feature Vectors

Feature Vectors

X names the collection of objects being studied, while Y names the collection of possible labels.

  • Programming

Two Questions Before Learning

A statistical learning problem begins by separating two questions: What objects are we studying, and what labels could describe those objects? The collection of objects is called the domain set and is written as X. The collection of possible labels is called the label set and is written as Y.

Consider a learning problem about papayas. The domain set X contains all papayas. The label set Y is {0, 1}, so the possible labels are 0 and 1. A papaya is an object being studied; 0 and 1 are possible labels for that object.

containscontainscontainsXall papayaspapayaone domain object0possible labelY{0, 1}1possible label
What does each set contain, and how are domain objects different from possible labels?

From Papaya to Features

An individual instance in X is called a domain point. A domain point is usually represented by a feature vector. For a papaya, color and softness are features. These features provide a representation of the papaya for the learning problem. They describe the domain point; they do not become labels.

A feature vector is a vector whose positions represent the features used to represent a domain point. In the papaya example, the relevant features are color and softness.

represented byfeature positionfeature positionpapayadomain pointfeature vectorvector representationposition 1colorposition 2softness
How does one domain object become a vector of feature values, and which feature corresponds to each vector position?

The Label Set

The label set Y is the collection of possible labels that may describe domain objects. In the papaya example, Y is {0, 1}. The values 0 and 1 are labels, while color and softness are features of a domain point.

containscontainsY{0, 1}0one possible class label1the other possible classlabel
What are the two possible values in the binary label set, and how do they represent the two classes?
CollectionNotationContains in the papaya exampleRole
Domain setXAll papayasObjects being studied
Label setY0 and 1Possible labels
Feature representationFeature vectorColor and softnessDescription of a domain point

Considering an Object and Label

Tracing One Papaya

Show how one papaya is handled in the learning setup.

Select the domain point: Begin with one particular papaya. It is an individual instance in X, so it belongs on the domain side of the problem.

Represent the domain point: Represent that papaya using a feature vector whose features include color and softness.

Consider a possible label: The possible labels come from Y, which is {0, 1} in this example. A possible label is therefore 0 or 1, not the papaya and not one of its features.

Form the labeled consideration: The domain point and a possible label are considered together as x from X and y from Y. The feature vector describes x, while y supplies the allowed label value.

The object remains a domain point in X, its representation uses features, and the possible label comes separately from Y.

represented bycombined withcombined withxone papaya in Xfeature vectorcolor and softness(x, y)object considered withlabely0 or 1 from Y
How are a domain object x from X and a possible label y from Y considered together as a labeled example?

The important conceptual trace is: object in X, feature representation for that object, and possible label from Y. The feature vector describes the domain point; the label set supplies the allowed label values.

Mistakes About X and Y

  • Putting a papaya in the label set

    A papaya is an object being studied, so it belongs on the domain side in X.

    Fix: Place all papayas, including one particular papaya, in X. Place the possible label values in Y.

  • Treating color and softness as labels

    Color and softness are features used to represent a domain point. The labels in the example are 0 and 1.

    Fix: Use color and softness in the feature representation, and use 0 or 1 as a possible label from Y.

  • Confusing a feature vector with the label set

    The feature vector describes a domain point, while Y is the collection of possible labels.

    Fix: Keep the representation of x separate from the possible label y.

Check the Separation

EASY

For the papaya learning problem, classify each item as belonging to X, Y, or the feature representation: all papayas, one particular papaya, 0, 1, color, and softness. Then describe the path from one papaya to a possible labeled example.

Hints
  • Start by asking whether the item is an object, a possible label, or a feature.
  • Remember that X contains the papayas and Y contains 0 and 1.
  • The path should include selecting a domain point, representing it with features, and considering a possible label from Y.
  1. X is the domain set: the collection of objects being studied. An individual object in X is a domain point. A domain point is usually represented by a feature vector; in the papaya example, color and softness are features. Y is the label set: the collection of possible labels. For papayas, Y is {0, 1}. A domain point and a possible label can be considered together, but the object, its features, and its label remain distinct concepts.

Key Takeaways

  • The domain set X contains the objects being studied.
  • An individual object in X is a domain point and is usually represented by a feature vector.
  • Color and softness are features of a papaya domain point, not labels.
  • The label set Y contains possible labels; for the papaya example, Y is {0, 1}.
  • A domain object x and a possible label y can be considered together while remaining distinct roles.