Feature Normalization
Feature manipulation transforms each original feature to create a resulting feature vector.
From Features to Vectors
A learning algorithm does not work directly with an abstract description of a problem. It works with a feature vector: a representation built from the original features. Feature manipulation changes those original features before the representation is given to the learning algorithm. Feature normalization is one kind of this manipulation.
The central idea is a transformation pipeline: original features are transformed, a resulting feature vector is created, and that vector is then used by the learning algorithm.
What Manipulation Is For
Feature manipulation is not merely a change in notation. Its purpose can be to reduce approximation or estimation errors, or to obtain a faster algorithm. These are different goals: one concerns the quality of the result, while the other concerns how efficiently the algorithm can operate.
| Possible purpose | What the transformation is intended to improve |
|---|---|
| Reduce approximation or estimation errors | The quality of the learned result |
| Obtain a faster algorithm | The speed of the learning procedure |
Feature manipulation can be motivated by accuracy-related or efficiency-related goals.
Choosing a Useful Transformation
There is no universally good feature transformation. The appropriate choice depends on the learning algorithm and on the prior assumptions about the problem. The same transformed representation may therefore be useful in one learning setting and less suitable in another.
Normalization in Linear Regression
Normalization is introduced as a type of feature manipulation through a linear regression problem using squared loss. In this setting, the data is described by a matrix whose rows are instance vectors and by a vector containing target values.
The important connection is representational. The rows of the matrix X are the instance vectors used by the regression problem. The vector y contains the target values. Normalization changes the feature representation that supplies those instance vectors; it does not replace the idea that the learning problem is described by instance vectors and target values.
A Vector-Level Example
Following One Instance Through a Transformation
Track an instance with two original features as feature manipulation creates a new representation.
Start with original features: Represent the instance by two original feature values, written as f1 and f2.
Transform each feature: Apply a feature transformation to each original feature, producing transformed values T1(f1) and T2(f2). The source concept does not require one universal transformation for every feature.
Build the resulting vector: Place the transformed values into a new feature vector: [T1(f1), T2(f2)].
Pass the representation to learning: The learning algorithm receives the resulting feature vector rather than the abstract problem description.
Feature manipulation can be understood as changing the representation of each instance before learning takes place.
This example deliberately leaves the transformation functions unspecified. The key lesson is the structure of the operation: original features are transformed, and the transformed values form the instance vector used by the algorithm.
Ridge Regression Connections
The normalization discussion also introduces ridge regression. In the stated linear regression setting, the rows of X are instance vectors and y is a vector of target values. Ridge regression returns a vector based on those instance vectors and target values.
Common Reasoning Errors
Treating normalization as universally beneficial
The appropriate transformation depends on the learning algorithm and the prior assumptions about the problem.
Fix:
Evaluate a transformation in the context of the algorithm and the problem assumptions.Confusing original features with the representation given to the algorithm
The algorithm works with a feature vector built from the original features.
Fix:
Describe the transformation from original features to the resulting feature vector.Reducing the purpose of manipulation to accuracy alone
The source identifies both reduced approximation or estimation errors and faster algorithms as possible purposes.
Fix:
Name the intended goal explicitly: error reduction, algorithmic speed, or both as applicable.Forgetting the roles of X and y
In the stated setting, the rows of X are instance vectors and y contains target values.
Fix:
Connect the transformed feature representation to the rows of X and connect the prediction targets to y.
Check Your Understanding
Explain the complete path from original features to a ridge-regression result. Your explanation should mention feature manipulation, the resulting feature vector, the rows of X, the target vector y, and the fact that the transformation must fit the learning algorithm and prior assumptions.
Hints
- Begin with what is changed before learning.
- State what the rows of X represent.
- State what y contains.
- End by explaining why the transformation cannot be selected in isolation.
What do you think happens?
If two feature transformations produce different resulting feature vectors, should you decide between them without considering the learning algorithm?
Reveal answer
Answer: No, because usefulness depends on the algorithm and prior assumptions.
Feature manipulation should fit the learning algorithm and the assumptions made about the problem. It is not universally good in isolation.
Key Takeaways
- Feature manipulation transforms original features to create a resulting feature vector.
- Its purposes can include reducing approximation or estimation errors and obtaining a faster algorithm.
- A useful transformation depends on the learning algorithm and prior assumptions about the problem.
- Normalization is motivated through linear regression with squared loss, where rows of X are instance vectors and y contains target values.
- Ridge regression returns a vector based on the instance vectors in X and the target values in y.
Key Takeaways
- Feature normalization is a form of feature manipulation that changes the representation supplied to a learning algorithm.
- Feature manipulation may target lower approximation or estimation errors or a faster algorithm.
- The right transformation must be judged together with the learning algorithm and prior assumptions about the problem.
- The normalization discussion uses linear regression with squared loss, instance vectors in the rows of X, and target values in y.
- Ridge regression operates on that representation and returns a vector based on X and y.