Feature Transformation
Feature learning learns a function ψ that maps instances from X into d-dimensional feature vectors.
From Instances to Features
A machine learning method cannot work directly with an unspecified instance space. It needs each instance to be expressed in a feature representation. Feature learning addresses this preparation step by learning a function ψ that maps an instance x from an instance space X into a d-dimensional feature vector ψ(x).
The important change occurs at the representation level. Before the mapping, an item is described only as an element of X. After the mapping, that item is available as a vector with d coordinates. The learned function creates the usable description on which a later learning method can operate.
Learning Versus Reworking Features
| Approach | Starting point | Main operation |
|---|---|---|
| Feature selection | A predefined feature space Rᵈ | Chooses some available features |
| Feature transformation | A predefined feature space Rᵈ | Changes individual features |
| Feature learning | An instance space X | Learns how instances should be represented |
Feature selection and feature transformation begin with a predefined vector space. Selection chooses from the features that already exist. Transformation changes those existing features. Feature learning instead learns a function that creates the representation from instances in X.
The No-Free-Lunch theorem implies that prior knowledge about the data distribution is necessary for building a good feature representation. In practice, this means that choosing or learning a representation still depends on understanding what properties of the data matter for the task.
Polynomial Features as a Mapping
Polynomial regression illustrates the general pattern clearly. Start with an input feature x. Construct a representation containing the monomials 1, x, x², through xᵖ. A linear predictor then operates on those constructed features.
Separating Representation from Prediction
Analyze the two stages in polynomial regression.
Construct features: The monomial construction maps the input into a representation containing 1, x, x², through xᵖ.
Apply the predictor: A linear predictor receives the constructed features and performs the later prediction task.
Identify the feature-learning idea: The transformation into the constructed representation is the part that illustrates feature learning; the later predictor is a separate part of the system.
Polynomial regression demonstrates that a representation can be constructed first and a linear predictor can be trained on top of it.
Location and Spread Adjustments
Some transformations reorganize the numerical position of a feature. Centering adjusts location, scaling adjusts range, and standardization adjusts both mean and variance. These operations change where values sit numerically rather than primarily changing how the feature's extreme values influence the representation.
| Transformation | Operation | Resulting adjustment |
|---|---|---|
| Centering | Subtract the empirical mean from every value | Changes the feature's location |
| Scaling | Adjust the feature's range | Commonly places values between 0 and 1 or between −1 and 1 |
| Standardization | Subtract the empirical mean and divide by the standard deviation | Produces zero mean and unit variance |
Centering changes the feature's location by subtracting its empirical mean from every value. Scaling changes the feature's range, commonly placing values between 0 and 1 or between −1 and 1. Standardization combines a location adjustment with a variance adjustment: it subtracts the empirical mean and divides by the standard deviation, producing a feature with zero mean and unit variance.
Reshaping Extreme Values
Other transformations do not primarily reposition a feature around a mean or place it inside a fixed range. They reshape how values, especially extreme values, influence the representation. Clipping, sigmoidal transformation, and logarithmic transformation address different patterns.
| Transformation | Useful when | Effect |
|---|---|---|
| Clipping | High or low values should be limited | Limits feature values to a specified range |
| Sigmoidal transformation | Extreme values should be softened rather than abruptly limited | Values close to zero change only slightly, while values far from zero behave similarly to clipping |
| Logarithmic transformation | A feature represents counts whose differences are not equally meaningful across the scale | Compresses the large-count end relative to the small-count end |
Clipping limits high or low feature values to a specified range. A sigmoidal transformation is a softer alternative: values close to zero are affected only slightly, while values far from zero behave similarly to clipping. A logarithmic transformation is especially useful for count features because equal numerical gaps may not carry equal meaning across the scale.
For a word-count feature, the difference between zero occurrences and one occurrence may matter much more than the difference between 1000 occurrences and 1001 occurrences. A logarithmic transformation compresses the large-count end relative to the small-count end, expressing those differences more suitably for the intended representation.
Choosing the Operation
The correct transformation depends on the property that needs to change. If the feature's location is the issue, centering addresses it. If its range is inconvenient, scaling is appropriate. If both mean and variance need adjustment, standardization combines those changes. If extreme values need a hard limit, use clipping; if they need a softer treatment, use a sigmoidal transformation. For count features in which small and large differences have unequal importance, a logarithmic transformation is useful.
Common Selection Mistakes
Treating feature learning as if it were only feature selection.
Feature selection starts with an existing feature space, whereas feature learning learns a function that creates the representation from instances.
Fix:
Ask whether the representation already exists. If it does, selection or transformation may be occurring; if the mapping from X is being learned, the process is feature learning.Using standardization as a synonym for every rescaling operation.
Scaling adjusts range, while standardization subtracts the empirical mean and divides by the standard deviation, producing zero mean and unit variance.
Fix:
Name the operation according to the property it changes: location, range, or both mean and variance.Choosing a transformation before identifying the behavior that needs to change.
Clipping limits values, while logarithmic transformation compresses the large-count end relative to the small-count end.
Fix:
Describe the desired change first, then select the transformation that directly addresses it.Confusing the constructed representation with the later predictor in polynomial regression.
The monomial construction defines the representation; the linear predictor operates after that representation has been created.
Fix:
Separate the mapping stage from the prediction stage.
Practice: Match the Property
For each situation, choose the transformation that most directly addresses the stated behavior: a feature is shifted away from zero; a feature must be placed in a convenient range; both mean and variance need adjustment; unusually high values must be limited; extreme values should be softened rather than abruptly limited; or a count feature gives unequal importance to gaps at the small and large ends.
Hints
- Centering changes location.
- Scaling changes range, while standardization changes mean and variance.
- Clipping is a hard limit, whereas a sigmoidal transformation is softer.
- Logarithmic transformation is especially useful for count features with unequal gap importance.
- Feature transformation is about changing how data is represented so that a learning method can use it more effectively. Some operations adjust numerical location and spread; others reshape the influence of extreme values or compress count scales.
Key Takeaways
- Feature learning learns a mapping ψ from an instance space X into d-dimensional feature vectors.
- Feature selection and feature transformation begin with a predefined feature space, while feature learning learns how instances should be represented.
- Polynomial regression separates feature construction from later prediction by mapping x into 1, x, x², through xᵖ before applying a linear predictor.
- Centering changes location, scaling changes range, and standardization produces zero mean and unit variance.
- Clipping, sigmoidal, and logarithmic transformations reshape the treatment of extreme values or count differences, so the desired change should guide the choice.