Concepts / Feature Selection

Feature Selection

Feature manipulation transforms each original feature to create a resulting feature vector.

  • Programming

From Features to a Vector

A learning algorithm does not work directly with an abstract description of a problem. It works with a feature vector: a representation built from the original features. Feature manipulation changes that representation before it is given to the learning algorithm. In this article, feature selection is treated through the source concept of feature manipulation: transforming each original feature to create a resulting feature vector.

transformtransformtransformFeature 1original featureVector position 1transformed featureFeature 2original featureVector position 2transformed featureFeature 3original featureVector position 3transformed feature
How does each original feature become part of the transformed feature vector, and how are the resulting feature positions related to the originals?

Tracing One Transformation

A Conceptual Feature Transformation

Suppose a problem begins with three original features. A feature manipulation is applied before learning.

Start with original features: The problem is described using Feature 1, Feature 2, and Feature 3.

Transform each feature: The feature manipulation changes the representation associated with each original feature.

Build the resulting vector: The transformed features occupy positions in a resulting feature vector.

Pass the representation to learning: The learning algorithm receives the resulting feature vector rather than the abstract problem description.

The important state change is from original features to a resulting feature vector. The transformation occurs before the learning algorithm uses the representation.

The example does not claim that one transformation is always correct. It shows the mechanism: original features are changed into a representation that the learning algorithm can use. The quality of that representation depends on the purpose of the manipulation, the selected learning algorithm, and the assumptions made about the problem.

Why Manipulate Features

Feature manipulation can serve at least two purposes. One purpose is to reduce approximation or estimation errors. Another is to obtain a faster algorithm. These are possible goals, not a guarantee that every transformation achieves both. A transformation should therefore be judged in relation to the learning problem rather than treated as universally beneficial.

may aim to reducemay aim to obtainFeaturemanipulationtransformed representationReduced errorsapproximation or estimationFaster algorithmcomputational purpose
How can transforming features reduce approximation or estimation error, and how can it make the learning algorithm faster?

Algorithm and Assumptions

A feature transformation is not universally good or bad. It is part of a larger design decision. First, the transformation produces the representation used by the learning algorithm. Second, the learning algorithm interprets that representation according to its own method. Third, prior assumptions about the problem influence whether the representation is appropriate. Changing the algorithm or changing those assumptions can therefore change which transformation is considered correct.

input toprepared forhelps determineusescreatesOriginal featuresFeaturetransformationLearning algorithmFeature vectorrepresentation used forlearningPrior assumptionsabout the problem
How does the same feature transformation interact with a chosen learning algorithm and with assumptions about the problem?

Normalization in Linear Regression

Normalization is introduced as a type of feature manipulation. Its motivation is presented through a linear regression problem using squared loss. In that setting, the data is described by a matrix whose rows are instance vectors and by a vector containing target values. Normalization changes the feature representation before that linear regression procedure uses it.

representation usedrepresentation usedOriginal featurevaluesbefore normalizationLinear regressionsquared lossNormalized featurevaluesafter normalizationLinear regressionsquared loss
How does changing the feature representation through normalization affect the setup of a linear regression problem with squared loss?

The before-and-after view is about the representation supplied to the model. It is not a claim that normalization has one universal numerical procedure or one universal effect. The source uses the linear regression and squared-loss setting to motivate normalization, so the learning algorithm and loss belong in the explanation.

Ridge Regression Data Flow

The normalization discussion also introduces ridge regression. In the stated linear regression setting, the rows of the matrix X are instance vectors, and y is a vector of target values. After feature normalization, the resulting instance vectors and their target values form the inputs to ridge regression. Ridge regression returns a vector based on those instance vectors and target values.

normalizeinstance vectorstarget valuesreturnsOriginal featuresNormalized instancevectorsrows of XRidge regressionCoefficient vectorreturned vectorTarget valuesvector y
How do normalized instance vectors and their target values flow into ridge regression to produce model coefficients?

Reading X and y Together

Interpret the roles of X and y in the stated ridge regression setup after feature normalization.

Identify X: X is the matrix whose rows are instance vectors. After normalization, those rows represent the transformed feature information used by the regression procedure.

Identify y: y is the vector containing target values associated with the instances.

Connect the inputs: Ridge regression uses the instance vectors and target values together.

Identify the result: Ridge regression returns a vector based on those instance vectors and target values.

Normalization changes the feature representation in X, while y supplies the target values for the stated ridge regression problem.

Mistakes to Avoid

  • Treating feature manipulation as something the learning algorithm does directly to an abstract problem description.

    The algorithm works with a feature vector, and feature manipulation changes that representation before the algorithm receives it.

    Fix: Trace the order explicitly: original features, transformation, resulting feature vector, then learning algorithm.

  • Assuming one feature transformation is best for every problem.

    The correct transformation depends on the learning algorithm and the prior assumptions about the problem.

    Fix: Evaluate the transformation together with the algorithm and the assumptions it is meant to support.

  • Claiming that feature manipulation always reduces error.

    Reducing approximation or estimation errors is one possible purpose, not a universal guarantee.

    Fix: Describe error reduction as a goal that must be considered in context.

  • Leaving target values out of the ridge regression setup.

    The stated setup includes both the rows of X, which are instance vectors, and y, which contains target values.

    Fix: Keep the normalized instance vectors and target values connected when describing the regression input.

Practice and Summary

MEDIUM

A problem begins with original features. You apply a transformation, obtain a resulting feature vector, and then use ridge regression. Explain what changed before learning, identify what the rows of X represent, identify what y represents, and name two possible purposes of the transformation.

Hints
  • Start with the distinction between original features and the resulting feature vector.
  • Recall the roles of X and y in the stated linear regression setting.
  • The two purposes concern errors and algorithm speed.
  1. Feature manipulation transforms original features into a resulting feature vector used by a learning algorithm. Its purposes may include reducing approximation or estimation errors and obtaining a faster algorithm. The appropriate transformation depends on the learning algorithm and prior assumptions about the problem. Normalization is presented as a feature manipulation motivated by linear regression with squared loss. In the associated ridge regression setup, the rows of X are instance vectors, y contains target values, and ridge regression returns a vector based on those inputs.

Key Takeaways

  • Feature manipulation transforms each original feature to create a resulting feature vector.
  • Its possible purposes include reducing approximation or estimation errors and obtaining a faster algorithm.
  • A transformation must be considered together with the learning algorithm and prior assumptions about the problem.
  • Normalization is introduced through linear regression with squared loss.
  • In the ridge regression setup, X contains instance vectors, y contains target values, and the procedure returns a vector based on them.