Mastering Feature Representation for Machine Learning
Feature transformations should be selected according to the property that needs to change.
Why Representation Matters
A feature is not always presented in the most useful numerical form. Its values may be shifted away from zero, cover an inconvenient range, contain unusually high or low values, or represent counts whose differences are not equally meaningful. Feature transformations change that representation. The key decision is not to transform automatically, but to identify which property needs to change first.
Following Location and Spread
Centering, scaling, and standardization reorganize the numerical position of a feature. They do not primarily change which observations are larger or smaller than one another. Instead, they alter where the values sit, how wide their range is, or how much variation the feature has around its center.
| Transformation | What changes | Source description |
|---|---|---|
| Centering | Feature location | Subtracts the empirical mean from every value |
| Scaling | Feature range | Commonly places values between 0 and 1 or between −1 and 1 |
| Standardization | Feature mean and variance | Subtracts the empirical mean and divides by the standard deviation |
The three location-and-spread transformations perform different jobs.
Choosing among centering, scaling, and standardization
A feature has values that are far from zero, span an inconvenient range, and vary by different amounts around their center. Which transformation matches each problem?
Values are shifted away from zero: Centering is the direct choice because it subtracts the empirical mean from every value and changes the feature's location.
The range is inconvenient: Scaling is the direct choice because it adjusts the feature's range, commonly to 0 through 1 or to −1 through 1.
Both location and variance need adjustment: Standardization combines a mean adjustment with division by the standard deviation, producing zero mean and unit variance.
Identify the property first: location suggests centering, range suggests scaling, and both mean and variance suggest standardization.
Reshaping Value Differences
Clipping, sigmoidal transformation, and logarithmic transformation address the influence or meaning of values rather than simply repositioning a feature around a mean. These operations reshape how differences are expressed. Two values can remain ordered in the same general direction while their numerical separation and influence change.
Clipping limits high or low values to a specified range. It directly prevents values from extending beyond those limits. A sigmoidal transformation is softer: values close to zero are affected only slightly, while values far from zero behave similarly to clipping. A logarithmic transformation is especially useful for count features because it compresses the large-count end relative to the small-count end.
Handling Extreme and Skewed Values
Selecting a reshaping operation
Three features have different problems: one contains unusually high and low values, one has extreme values that should be softened rather than abruptly limited, and one records counts where small and large differences do not have equal importance.
Unusually high or low values must stay within explicit limits: Use clipping because it limits high or low feature values to a specified range.
Extreme values need a softer treatment: Use a sigmoidal transformation because values near zero are affected only slightly while far-from-zero values behave similarly to clipping.
The feature represents counts: Use a logarithmic transformation when small and large count differences have unequal importance; it compresses the large-count end relative to the small-count end.
The observed behavior determines the choice: hard limits suggest clipping, softer extreme compression suggests sigmoid, and unequal importance across a count scale suggests logarithm.
Choosing by Feature Behavior
This selection process can also lead to a combination. For example, centering and scaling may be used together when both the feature's location and range need adjustment. The important discipline is to identify the desired property first and then choose the operation or combination that addresses it.
Treating every transformation as if it did the same job
Scaling adjusts range, while clipping addresses high or low values by limiting them to a specified range.
Fix:
Name the behavior that needs to change before selecting the transformation.Confusing centering with standardization
Centering changes location only. Standardization subtracts the empirical mean and divides by the standard deviation.
Fix:
Use standardization when both mean and variance need adjustment.Using clipping whenever extreme values appear
A sigmoidal transformation provides a softer alternative: near-zero values are affected only slightly and extreme values are compressed.
Fix:
Choose clipping for specified limits and sigmoid for softer extreme-value behavior.Treating all count differences as equally meaningful
For count features, small and large differences may have unequal importance.
Fix:
Consider a logarithmic transformation when the large-count end should be compressed relative to the small-count end.
Practice the Selection
For each situation, choose the most appropriate transformation and explain which feature property it changes: a feature is shifted away from zero; a feature covers an inconvenient range; a feature needs both its mean and variance adjusted; high and low values must be limited; extreme values should be compressed softly; or a count feature has unequal meaning across small and large values.
Hints
- Centering addresses location.
- Scaling addresses range.
- Standardization addresses mean and variance together.
- Clipping uses specified limits, while sigmoid provides a softer treatment of extremes.
- Logarithm compresses the large-count end relative to the small-count end.
What do you think happens?
A feature records word counts. Moving from zero occurrences to one occurrence is more meaningful than moving from 1000 occurrences to 1001 occurrences. Which transformation best matches this behavior?
Reveal answer
Answer: Logarithmic transformation
A logarithmic transformation compresses the large-count end relative to the small-count end, matching a count feature whose differences do not have equal importance across the scale.
Practical Takeaways
- Centering changes a feature's location by subtracting its empirical mean from every value.
- Scaling changes a feature's range, commonly placing it between 0 and 1 or between −1 and 1.
- Standardization adjusts both mean and variance by subtracting the empirical mean and dividing by the standard deviation.
- Clipping and sigmoid both address extreme values, but clipping imposes specified limits while sigmoid provides softer compression.
- Logarithmic transformation is useful for count features when small and large differences have unequal importance.
Key Takeaways
- Select a feature transformation according to the property that needs to change.
- Use centering for location, scaling for range, and standardization for both mean and variance.
- Use clipping for explicit limits and sigmoid for softer treatment of extreme values.
- Use logarithmic transformation when a count feature gives unequal importance to small and large differences.
- Transformations can be combined, but the desired change should be identified first.