Polynomial Regression
Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.
From an Instance to a Representation
A machine learning method cannot work directly with an unspecified instance space. It needs each instance to be expressed in a usable feature representation. Feature learning addresses this preparation step by learning how an instance should be represented instead of requiring the representation to be completely designed in advance.
Feature learning learns a function ψ that maps instances from an instance space X into d-dimensional feature vectors.
The important change occurs at the representation level. Before the mapping, an item is described only as an element of X. After the mapping, that same item is available as a vector with d coordinates. The function ψ creates the description that a later learning method can use.
Polynomial Terms as Constructed Features
Polynomial regression illustrates the general pattern of constructing features first and training a linear predictor on top of them. For an input x, the constructed feature vector can contain the monomial terms 1, x, x², through xᵈ. The representation is created before the later predictor operates on it.
Constructing a Degree-3 Representation
Represent the input x = 2 using the monomial construction through degree 3.
Start with the input: The instance supplied to the representation step is x = 2.
Construct the terms: The terms through degree 3 are 1, x, x², and x³.
Evaluate the terms: For x = 2, these terms are 1, 2, 4, and 8.
Separate representation from prediction: The resulting vector is the constructed representation. A later linear predictor is a separate part of the learning system.
The constructed feature vector is [1, 2, 4, 8].
This example separates two roles that are easy to confuse. The monomial construction defines a representation, while the linear predictor operates after that representation has been created. Feature learning concerns the transformation into the representation; the later predictor is a separate part of the learning system.
Degree and Feature-Space Size
The degree controls which monomial terms are included in the constructed representation. Through degree d, the vector is described as [1, x, x², ..., xᵈ]. Increasing the degree therefore adds higher-order constructed terms to the representation. The meaning of the feature space changes because the later predictor receives a richer collection of coordinates.
Learning versus Choosing Features
Feature selection and feature transformation begin with a predefined feature space, described in the source as Rᵈ. Feature selection chooses some of the available features. Feature transformation changes individual features. Feature learning starts one step earlier: it learns a function that creates the representation from instances in X.
| Approach | Starting point | Main operation |
|---|---|---|
| Feature learning | Instances in X | Learn a function that creates the representation |
| Feature selection | Predefined feature space Rᵈ | Choose some available features |
| Feature transformation | Predefined feature space Rᵈ | Change individual features |
Why Prior Knowledge Matters
The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation. A representation is not useful merely because it is automated. Its usefulness is connected to whether it produces a suitable hypothesis class for the task and the distribution of the data.
Common Conceptual Mistakes
Calling polynomial regression only a linear-predictor procedure
Polynomial regression illustrates a two-part pattern: first construct features, then train a linear predictor on those features.
Fix:
Separate the monomial construction from the predictor that operates after the representation exists.Treating feature learning as ordinary feature selection
Feature selection begins with a predefined feature space, whereas feature learning starts with instances in X and learns how to create their representation.
Fix:
Ask whether the representation already exists before the method acts. If it does not, the task is at the feature-learning level.Assuming automation eliminates the need for prior knowledge
The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation.
Fix:
Connect the representation to the data distribution and the hypothesis class needed for the task.
Check Your Understanding
An instance belongs to X, and a method learns a function ψ that maps it into a vector with d coordinates. Is this feature selection, feature transformation, or feature learning? Explain what happens before and after the mapping.
Hints
- Identify whether the feature space was predefined before the method acted.
- Use the distinction between an instance in X and a vector in a predefined Rᵈ.
For x = 3, write the constructed monomial feature vector through degree 2. Then state which part of the polynomial-regression pattern is represented by that vector and which part is still separate.
Hints
- List the terms 1, x, and x².
- The vector is the representation; the later linear predictor is separate.
What do you think happens?
If the polynomial degree increases from 2 to 3, what changes in the constructed representation?
Reveal answer
Answer: A higher-order term is added.
The representation through degree 2 contains 1, x, and x². Through degree 3, it additionally contains x³.
Key Takeaways
- Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.
- Polynomial regression illustrates feature construction followed by a linear predictor.
- The monomial construction creates the representation; the later predictor is a separate system component.
- Feature selection and feature transformation begin with a predefined feature space, while feature learning starts by learning how instances should be represented.
- Prior knowledge about the data distribution is necessary for constructing a good feature representation.
Key Takeaways
- Feature learning maps an instance from X into a d-dimensional feature vector through a learned function ψ.
- Polynomial regression demonstrates how constructed features such as 1, x, x², through xᵈ can be created before prediction.
- The feature-construction step and the later linear predictor play different roles.
- Feature selection and transformation operate on a predefined feature space, whereas feature learning learns the representation itself.
- Prior knowledge about the data distribution remains necessary for choosing or learning a useful representation.