Concepts / Polynomial Regression

Polynomial Regression

Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.

  • Programming

From an Instance to a Representation

A machine learning method cannot work directly with an unspecified instance space. It needs each instance to be expressed in a usable feature representation. Feature learning addresses this preparation step by learning how an instance should be represented instead of requiring the representation to be completely designed in advance.

Feature learning learns a function ψ that maps instances from an instance space X into d-dimensional feature vectors.

inputmaps toInstance xelement of XFunction ψlearned representationmappingFeature vectord coordinates
What happens to an instance x when the feature-learning function ψ maps it from the original instance space X into a d-dimensional feature vector?

The important change occurs at the representation level. Before the mapping, an item is described only as an element of X. After the mapping, that same item is available as a vector with d coordinates. The function ψ creates the description that a later learning method can use.

Polynomial Terms as Constructed Features

Polynomial regression illustrates the general pattern of constructing features first and training a linear predictor on top of them. For an input x, the constructed feature vector can contain the monomial terms 1, x, x², through xᵈ. The representation is created before the later predictor operates on it.

provide xconstructssupplies featuresInput xinstanceMonomial constructionrepresentation stepFeature vector[1, x, x², ..., xᵈ]Linear predictoroperates after construction
How does an input x become the constructed feature vector [1, x, x², ..., xᵈ] used by polynomial regression?

Constructing a Degree-3 Representation

Represent the input x = 2 using the monomial construction through degree 3.

Start with the input: The instance supplied to the representation step is x = 2.

Construct the terms: The terms through degree 3 are 1, x, x², and x³.

Evaluate the terms: For x = 2, these terms are 1, 2, 4, and 8.

Separate representation from prediction: The resulting vector is the constructed representation. A later linear predictor is a separate part of the learning system.

The constructed feature vector is [1, 2, 4, 8].

This example separates two roles that are easy to confuse. The monomial construction defines a representation, while the linear predictor operates after that representation has been created. Feature learning concerns the transformation into the representation; the later predictor is a separate part of the learning system.

Degree and Feature-Space Size

The degree controls which monomial terms are included in the constructed representation. Through degree d, the vector is described as [1, x, x², ..., xᵈ]. Increasing the degree therefore adds higher-order constructed terms to the representation. The meaning of the feature space changes because the later predictor receives a richer collection of coordinates.

degree 1degree 3xdegree 1 inputxdegree 3 input[1, x]constant and first-orderterms[1, x, x², x³]constant, first-, second-,and third-order terms
How does increasing the polynomial degree change the number and meaning of the constructed features?

Learning versus Choosing Features

Feature selection and feature transformation begin with a predefined feature space, described in the source as Rᵈ. Feature selection chooses some of the available features. Feature transformation changes individual features. Feature learning starts one step earlier: it learns a function that creates the representation from instances in X.

learn ψchoosechangeInstance in Xrepresentation not fixed inadvancePredefined Rᵈavailable feature spacePredefined featurefeature already availableLearned vectorcreated by ψSelected featuressubset of availablefeaturesChanged featuretransformed existingfeature
What is the difference between learning a mapping ψ and selecting or transforming features that were predefined in advance?
ApproachStarting pointMain operation
Feature learningInstances in XLearn a function that creates the representation
Feature selectionPredefined feature space RᵈChoose some available features
Feature transformationPredefined feature space RᵈChange individual features

Why Prior Knowledge Matters

The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation. A representation is not useful merely because it is automated. Its usefulness is connected to whether it produces a suitable hypothesis class for the task and the distribution of the data.

informsproduceslearnssupportsPrior knowledgeabout data distributionFeaturerepresentationmapping ψHypothesis classsuitable for the taskAutomationlearns the transformationLearning tasktarget use
How does prior knowledge about the data distribution influence which features or polynomial terms are included in a representation?

Common Conceptual Mistakes

  • Calling polynomial regression only a linear-predictor procedure

    Polynomial regression illustrates a two-part pattern: first construct features, then train a linear predictor on those features.

    Fix: Separate the monomial construction from the predictor that operates after the representation exists.

  • Treating feature learning as ordinary feature selection

    Feature selection begins with a predefined feature space, whereas feature learning starts with instances in X and learns how to create their representation.

    Fix: Ask whether the representation already exists before the method acts. If it does not, the task is at the feature-learning level.

  • Assuming automation eliminates the need for prior knowledge

    The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation.

    Fix: Connect the representation to the data distribution and the hypothesis class needed for the task.

Check Your Understanding

EASY

An instance belongs to X, and a method learns a function ψ that maps it into a vector with d coordinates. Is this feature selection, feature transformation, or feature learning? Explain what happens before and after the mapping.

Hints
  • Identify whether the feature space was predefined before the method acted.
  • Use the distinction between an instance in X and a vector in a predefined Rᵈ.
MEDIUM

For x = 3, write the constructed monomial feature vector through degree 2. Then state which part of the polynomial-regression pattern is represented by that vector and which part is still separate.

Hints
  • List the terms 1, x, and x².
  • The vector is the representation; the later linear predictor is separate.

What do you think happens?

If the polynomial degree increases from 2 to 3, what changes in the constructed representation?

  • Nothing changes
  • A higher-order term is added
  • The instance space X disappears
  • Feature selection replaces feature learning
Reveal answer

Answer: A higher-order term is added.

The representation through degree 2 contains 1, x, and x². Through degree 3, it additionally contains x³.

Key Takeaways

  1. Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.
  2. Polynomial regression illustrates feature construction followed by a linear predictor.
  3. The monomial construction creates the representation; the later predictor is a separate system component.
  4. Feature selection and feature transformation begin with a predefined feature space, while feature learning starts by learning how instances should be represented.
  5. Prior knowledge about the data distribution is necessary for constructing a good feature representation.

Key Takeaways

  • Feature learning maps an instance from X into a d-dimensional feature vector through a learned function ψ.
  • Polynomial regression demonstrates how constructed features such as 1, x, x², through xᵈ can be created before prediction.
  • The feature-construction step and the later linear predictor play different roles.
  • Feature selection and transformation operate on a predefined feature space, whereas feature learning learns the representation itself.
  • Prior knowledge about the data distribution remains necessary for choosing or learning a useful representation.