Concepts / Feature Transformation

Feature Transformation

Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.

  • Programming

From Instances to Features

A machine learning method cannot work directly with an unspecified instance space. It needs each instance to be expressed in a feature representation. Feature learning addresses this preparation step by learning a function ψ that maps an instance x from an instance space X into a d-dimensional feature vector ψ(x).

mapped bymapped byproducesproducesx₁instance in Xψlearned mappingψ(x₁)d-dimensional vectorx₂instance in Xψ(x₂)d-dimensional vector
How does an instance x in the original space X become a d-dimensional feature vector ψ(x), and how do multiple instances map into the resulting feature space?

The important change occurs at the representation level. Before the mapping, an item is described only as an element of X. After the mapping, that item is available as a vector with d coordinates. The learned function creates the usable description on which a later learning method can operate.

Learning Versus Reworking Features

ApproachStarting pointMain operation
Feature selectionA predefined feature space RᵈChooses some available features
Feature transformationA predefined feature space RᵈChanges individual features
Feature learningAn instance space XLearns how instances should be represented

Feature selection and feature transformation begin with a predefined vector space. Selection chooses from the features that already exist. Transformation changes those existing features. Feature learning instead learns a function that creates the representation from instances in X.

The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation. In practice, this means that choosing or learning a representation still depends on understanding what properties of the data matter for the task.

Polynomial Features as a Mapping

Polynomial regression illustrates the general pattern clearly. Start with an input feature x. Construct a representation containing the monomials 1, x, x², through xᵖ. A linear predictor then operates on those constructed features.

transformconstructsconstructsconstructsconstructsinputinputinputinputxinput featureψ(x)monomial construction1constructed featureLinear predictoroperates after mappingxconstructed featurex²constructed featurexᵖconstructed feature
How does a single input feature x become the constructed features 1, x, x², ..., xᵖ before regression combines them?

Separating Representation from Prediction

Analyze the two stages in polynomial regression.

Construct features: The monomial construction maps the input into a representation containing 1, x, x², through xᵖ.

Apply the predictor: A linear predictor receives the constructed features and performs the later prediction task.

Identify the feature-learning idea: The transformation into the constructed representation is the part that illustrates feature learning; the later predictor is a separate part of the system.

Polynomial regression demonstrates that a representation can be constructed first and a linear predictor can be trained on top of it.

Location and Spread Adjustments

Some transformations reorganize the numerical position of a feature. Centering adjusts location, scaling adjusts range, and standardization adjusts both mean and variance. These operations change where values sit numerically rather than primarily changing how the feature's extreme values influence the representation.

subtract meanadjust rangesubtract mean and divide by standard deviationOriginal featureoriginal location andspreadCentered featuresubtract empirical meanScaled featureadjust rangeStandardized featuremean and variance adjusted
What changes in a feature's values, center, and spread when we subtract the mean, divide by a scale, or standardize it?
TransformationOperationResulting adjustment
CenteringSubtract the empirical mean from every valueChanges the feature's location
ScalingAdjust the feature's rangeCommonly places values between 0 and 1 or between −1 and 1
StandardizationSubtract the empirical mean and divide by the standard deviationProduces zero mean and unit variance

Centering changes the feature's location by subtracting its empirical mean from every value. Scaling changes the feature's range, commonly placing values between 0 and 1 or between −1 and 1. Standardization combines a location adjustment with a variance adjustment: it subtracts the empirical mean and divides by the standard deviation, producing a feature with zero mean and unit variance.

Reshaping Extreme Values

Other transformations do not primarily reposition a feature around a mean or place it inside a fixed range. They reshape how values, especially extreme values, influence the representation. Clipping, sigmoidal transformation, and logarithmic transformation address different patterns.

producesproducesproducesClippinglimits values to aspecified rangeSpecified rangehigh and low values limitedSigmoidaltransformationsoft treatment of extremesSimilar extremebehaviorvalues far from zerotreated similarlyLogarithmictransformationcompresses large-countdifferencesCompressed largecountssmall-count differencesemphasized relatively
How do clipping, sigmoid, and logarithmic functions change extreme values, boundedness, and the relative spacing of feature values?
TransformationUseful whenEffect
ClippingHigh or low values should be limitedLimits feature values to a specified range
Sigmoidal transformationExtreme values should be softened rather than abruptly limitedValues close to zero change only slightly, while values far from zero behave similarly to clipping
Logarithmic transformationA feature represents counts whose differences are not equally meaningful across the scaleCompresses the large-count end relative to the small-count end

Clipping limits high or low feature values to a specified range. A sigmoidal transformation is a softer alternative: values close to zero are affected only slightly, while values far from zero behave similarly to clipping. A logarithmic transformation is especially useful for count features because equal numerical gaps may not carry equal meaning across the scale.

For a word-count feature, the difference between zero occurrences and one occurrence may matter much more than the difference between 1000 occurrences and 1001 occurrences. A logarithmic transformation compresses the large-count end relative to the small-count end, expressing those differences more suitably for the intended representation.

Choosing the Operation

checkcheckcheckcheckcheckchoosechoosechoosehard limitsoft limitchooseFeature behavioridentify the property tochangeLocationvalues shifted from desiredcenterCenteringsubtract empirical meanRangevalues span an inconvenientintervalScalingadjust rangeMean and varianceboth need adjustmentStandardizationzero mean and unit varianceExtreme valueshigh or low values dominateClippinglimit to specified rangeUnequal count meaningsmall and large gaps differin importanceSigmoidaltransformationsoften extremesLogarithmictransformationcompress large counts
How does the observed behavior of a feature, such as outliers, skew, saturation, or an expected range, lead to selecting one transformation rather than another?

The correct transformation depends on the property that needs to change. If the feature's location is the issue, centering addresses it. If its range is inconvenient, scaling is appropriate. If both mean and variance need adjustment, standardization combines those changes. If extreme values need a hard limit, use clipping; if they need a softer treatment, use a sigmoidal transformation. For count features in which small and large differences have unequal importance, a logarithmic transformation is useful.

Common Selection Mistakes

  • Treating feature learning as if it were only feature selection.

    Feature selection starts with an existing feature space, whereas feature learning learns a function that creates the representation from instances.

    Fix: Ask whether the representation already exists. If it does, selection or transformation may be occurring; if the mapping from X is being learned, the process is feature learning.

  • Using standardization as a synonym for every rescaling operation.

    Scaling adjusts range, while standardization subtracts the empirical mean and divides by the standard deviation, producing zero mean and unit variance.

    Fix: Name the operation according to the property it changes: location, range, or both mean and variance.

  • Choosing a transformation before identifying the behavior that needs to change.

    Clipping limits values, while logarithmic transformation compresses the large-count end relative to the small-count end.

    Fix: Describe the desired change first, then select the transformation that directly addresses it.

  • Confusing the constructed representation with the later predictor in polynomial regression.

    The monomial construction defines the representation; the linear predictor operates after that representation has been created.

    Fix: Separate the mapping stage from the prediction stage.

Practice: Match the Property

MEDIUM

For each situation, choose the transformation that most directly addresses the stated behavior: a feature is shifted away from zero; a feature must be placed in a convenient range; both mean and variance need adjustment; unusually high values must be limited; extreme values should be softened rather than abruptly limited; or a count feature gives unequal importance to gaps at the small and large ends.

Hints
  • Centering changes location.
  • Scaling changes range, while standardization changes mean and variance.
  • Clipping is a hard limit, whereas a sigmoidal transformation is softer.
  • Logarithmic transformation is especially useful for count features with unequal gap importance.
  1. Feature transformation is about changing how data is represented so that a learning method can use it more effectively. Some operations adjust numerical location and spread; others reshape the influence of extreme values or compress count scales.

Key Takeaways

  • Feature learning learns a mapping ψ from an instance space X into d-dimensional feature vectors.
  • Feature selection and feature transformation begin with a predefined feature space, while feature learning learns how instances should be represented.
  • Polynomial regression separates feature construction from later prediction by mapping x into 1, x, x², through xᵖ before applying a linear predictor.
  • Centering changes location, scaling changes range, and standardization produces zero mean and unit variance.
  • Clipping, sigmoidal, and logarithmic transformations reshape the treatment of extreme values or count differences, so the desired change should guide the choice.