Linear Predictors
Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.
From Instances to Usable Representations
A machine learning method cannot work directly with an unspecified instance space. It needs each instance to be expressed in a feature representation. Feature learning addresses this preparation step by learning a function ψ that maps an instance from X into a d-dimensional feature vector.
The important change occurs at the representation level. Before the mapping, an item is described only as an element of X. After the mapping, that item is available as a vector with d coordinates. Feature learning is therefore the process of learning the transformation ψ that creates this usable description.
Three Ways to Obtain Features
| Approach | Starting point | What is learned or done |
|---|---|---|
| Feature learning | Instances in X | A function that creates the representation |
| Feature selection | A predefined feature space Rᵈ | Some available features are selected |
| Feature transformation | A predefined feature space Rᵈ | Individual features are changed |
Polynomial Features as a Construction Pattern
Polynomial regression illustrates the general pattern of constructing features first and training a linear predictor on top of them. An input x can be mapped into a constructed feature vector containing terms such as 1, x, x², and higher powers. The monomial construction defines the representation; the later linear predictor operates after that representation has been created.
Separating Representation from Prediction
Consider the polynomial feature construction that maps an input x into the feature vector containing 1, x, and x².
Construct the representation: The mapping creates the coordinates 1, x, and x². This construction is the representation step.
Apply the predictor: A linear predictor can then operate on the constructed vector. It is a separate step from creating the polynomial features.
Keep the roles distinct: The monomial construction defines how the instance is represented, while the later predictor uses that representation.
Polynomial regression demonstrates the pattern of constructing features first and training a linear predictor on top of them.
Projecting Instances with a Shared Vector
A linear ranking function applies the same vector w to instances through f(x) = w · x. For each instance, the projection produces one scalar. The vector w is shared across evaluations, while the instance x changes from one evaluation to the next.
The scalar is the result of evaluating the ranking function on one instance. It is not the entire feature vector. The vector x supplies the instance-specific input, and w supplies the projection rule that is reused for every instance.
From Scores to an Ordering
Three Instances, Three Scalar Outputs
Use the generated vectors w = (2, 1), x₁ = (3, 4), x₂ = (1, 2), and x₃ = (4, 0) to illustrate how one shared vector produces scalar outputs.
Evaluate x₁: Applying the same vector w to x₁ gives the scalar 2 × 3 + 1 × 4 = 10.
Evaluate x₂: Applying w to x₂ gives the scalar 2 × 1 + 1 × 2 = 4.
Evaluate x₃: Applying w to x₃ gives the scalar 2 × 4 + 1 × 0 = 8.
Compare the outputs: The instances have become the scalar outputs 10, 4, and 8 under the same projection rule.
Sorting the scalar outputs from largest to smallest gives the ranking x₁, x₃, x₂.
This example separates two ideas. First, the projection computes a scalar for each instance. Second, the collection of scalar results represents the ranking function and can be compared to form an ordering. Changing x changes the evaluated output; changing w changes the projection rule used for every evaluation.
Mistakes in Representation and Ranking
Treating feature selection as feature learning
Feature learning starts one step earlier by learning how instances in X should be represented.
Fix:
Ask whether the procedure begins with a predefined feature space or learns a mapping from instances into a feature vector space.Confusing the representation with the predictor
The construction defines the feature representation, while the predictor operates after that representation exists.
Fix:
Trace the two stages separately: first construct ψ(x), then apply the predictor.Thinking that one scalar is the whole representation of every instance
The feature vector describes the instance for the predictor; the scalar is the output of applying the ranking function to that instance.
Fix:
Label the stages explicitly as feature vector input, projection, and scalar output.Forgetting that the same w is used across instances
The ranking function applies the same vector w to instances, while x changes from one evaluation to the next.
Fix:
Hold w fixed while evaluating each instance, then compare the resulting scalars.
Apply the Two-Stage Trace
A system begins with instances in X. It first learns a mapping ψ into d-dimensional vectors and then applies the same vector w to each resulting vector. Explain what changes when the instance changes, what remains fixed during the ranking evaluations, and why the scalar outputs can be used to order the instances.
Hints
- Separate the representation stage from the prediction stage.
- Identify the object that changes from one evaluation to the next.
- Identify the vector that is shared across evaluations.
- The final comparison is between scalar outputs.
- Linear predictors become easier to understand when the process is traced in order: an instance begins in X, a feature-learning map ψ creates a d-dimensional vector, a shared vector w is applied to that representation, and the resulting scalar is compared with the scalars from other instances.
Key Takeaways
- Feature learning learns a mapping ψ from instances in X into d-dimensional feature vectors.
- Feature selection and feature transformation begin with a predefined feature space, whereas feature learning learns how the representation should be created.
- Polynomial regression illustrates constructing features first and applying a linear predictor afterward.
- A linear ranking function applies the same vector w to each instance and produces one scalar per instance.
- Comparing and sorting those scalar outputs provides an ordering of the instances.