Predictor design
Feature selection reduces a full feature collection to a subset used by the predictor.
A smaller input can be better
A machine-learning instance can be represented by many features, but a useful predictor may need only a small portion of them. Feature selection is the process of choosing that smaller, relevant subset for use in the predictor. The goal is not simply to remove information. It is to design a predictor that relies on a limited collection of features while still supporting useful prediction.
Mapping the selected inputs
Think of feature selection as a mapping from a full feature collection to the smaller collection that the predictor actually receives. Features that are not selected do not become predictor inputs. The retained features keep their role as meaningful inputs, but the predictor now works with fewer of them.
Selecting inputs for a hypothetical predictor
A predictor begins with six available features. A feature-selection procedure retains three features because they form a useful smaller input set.
Start with the full collection: The predictor could potentially use all six available features.
Choose a subset: The design retains three features and excludes the other three from the predictor's inputs.
Use the reduced inputs: The resulting predictor stores and processes the selected three-feature collection rather than the full collection.
The predictor has a smaller input collection. This can reduce storage and processing requirements, provided the retained features still support useful prediction.
Why fewer features matter
The smaller subset changes the predictor in two direct ways. First, there are fewer inputs to store. This can lower memory use. Second, there are fewer inputs to process when the predictor is applied, which can make application faster. Feature acquisition can also have a cost, so a predictor that needs only a small number of test results may be preferable in a medical-diagnosis setting, even if using fewer features causes a small drop in performance compared with using more features.
In a medical-diagnosis setting, possible features can correspond to test results. A predictor that needs only a small number of test results may reduce the cost of obtaining those features. That benefit can matter even when the smaller input set produces a small performance drop compared with using more features.
Estimation error and overfitting
Using fewer features can affect statistical behavior as well as operating cost. Restricting the predictor to a small subset can reduce its estimation error and therefore help prevent overfitting. In practical terms, the predictor is not given every available input merely because those inputs exist. It is limited to a collection chosen for useful prediction.
Why exhaustive search breaks down
The most direct strategy is exhaustive search. For every candidate subset of the available features, we could consider a predictor and then choose the best-performing candidate. This is conceptually attractive because it appears to examine every possible choice.
As the feature collection becomes larger, the number of candidate subsets that exhaustive search must consider grows until the search becomes computationally infeasible in typical situations. Exhaustive search is therefore an ideal description of what we would like to do, rather than a generally practical procedure.
Practical search strategies
| Approach | Subsets considered | Practical implication |
|---|---|---|
| Exhaustive search | Every possible candidate subset | Conceptually direct, but usually computationally infeasible as the feature collection grows |
| Computationally feasible selection | A limited search chosen by the method | Avoids the full search and seeks a subset that performs reasonably well |
Computationally feasible feature-selection approaches accept that the selected subset may not be optimal in exchange for a procedure that works reasonably well in practice. The source introduces three such approaches without detailing their individual procedures here. Their shared purpose is to find a useful subset without requiring the full exhaustive search.
Mistakes in feature selection
Treating feature selection as removing information without regard to prediction
The selected subset must still support useful prediction. A smaller input set can cause a performance drop if relevant information is removed.
Fix:
Evaluate feature selection as a trade-off between a limited input collection and useful predictive performance.Assuming that testing every subset is the normal practical solution
The number of candidate subsets grows as the feature collection becomes larger, making exhaustive search computationally infeasible in typical situations.
Fix:
Use a computationally feasible approach when a full exhaustive search is not practical.Confusing a feasible subset with a guaranteed optimal subset
Feasible methods accept that the selected subset may not be optimal in exchange for a procedure that works reasonably well in practice.
Fix:
Recognize and evaluate the practical compromise between search cost and subset quality.Considering only storage and processing costs
The cost of acquiring a feature can matter as well as the cost of storing or processing it.
Fix:
Include feature-acquisition cost when judging whether a smaller predictor design is preferable.
Check your design choice
A predictor can use a full collection of features or a smaller selected subset. Explain why the smaller design might be preferred, and then explain why selecting the smallest possible subset is not automatically the correct decision. Finally, compare exhaustive search with a computationally feasible selection approach.
Hints
- Consider memory use, prediction time, feature-acquisition cost, estimation error, and overfitting.
- Exhaustive search evaluates every possible subset; feasible approaches avoid requiring the full search.
- The selected subset should remain useful for prediction.
- Feature selection reduces a full feature collection to a relevant subset used by a predictor. Fewer inputs can reduce memory use, prediction time, and feature-acquisition cost. Restricting the input set can also reduce estimation error and help prevent overfitting, although removing too much information can hurt prediction. Exhaustive search examines every possible subset but is usually computationally infeasible as the feature collection grows. Practical selection methods search less, accept that their subset may not be optimal, and aim for a useful result that can be obtained in practice.
Key Takeaways
- Feature selection chooses a relevant subset from a full feature collection for use by a predictor.
- Fewer features can reduce memory use, prediction time, and the cost of acquiring inputs.
- A restricted input set can reduce estimation error and help prevent overfitting, but the subset must still support useful prediction.
- Exhaustive search considers every possible subset and usually becomes computationally infeasible as the feature collection grows.
- Computationally feasible methods trade possible optimality for a selection procedure that works reasonably well in practice.