Predictive Models and Quality Measures
Filter methods reduce a feature set by scoring each feature independently.
Why Features Need Individual Scores
Suppose a dataset contains many candidate features, but a later predictive model should use only a smaller set. A filter method addresses this by examining each feature on its own, assigning the feature a quality score, and then using the scores to choose candidates. The important idea is that assessment happens before the features are combined in a larger predictive model.
A filter method separates two decisions. First, it assesses each feature independently. Second, it compares the resulting scores and selects features. This turns one large feature-selection question into a collection of single-feature quality judgments.
Tracing One Feature Through Assessment
To score a particular feature, the filter method takes the values of that feature across the training examples and pairs them with the matching target values. The feature under assessment is the only input considered by the predictor used for that assessment.
A Single-Feature Quality Judgment
A dataset has several candidate features. The filter method is currently assessing Feature B using an error-based predictive measure.
Isolate the feature: Take the values of Feature B across the training examples. Do not combine Feature B with the other candidate features during this assessment.
Pair values with targets: Use the Feature B values together with the matching target values from those same training examples.
Train the one-feature predictor: Train a predictor using only Feature B. The resulting predictor is used to judge how well this feature supports predictions.
Measure prediction error: Compare the predictor's predictions with the matching target values using the selected quality measure.
Record one result: The assessment produces one quality score for Feature B. The other features will receive their own scores through their own single-feature assessments.
Feature B receives an individual quality score that can later be compared with the scores of the other features.
Choosing a Quality Measure
A filter method does not prescribe one universal way to judge a feature. The quality measure supplies that judgment, and many quality measures are possible. One direct approach is predictive: train a predictor using only the feature under assessment, then use the predictor's error as the feature's score.
An error-based feature score is a quality measure obtained from the prediction error of a predictor trained solely from the feature being assessed.
For the regression illustration, the feature values and their matching target values define the single-feature evaluation. The predictor is not trained from the whole candidate feature set for this step. It is trained from the one feature whose quality is being measured.
Reading Empirical Squared Loss
The source illustrates the error-based approach with linear regression and squared loss. For one feature, the predictor produces a prediction for each training example. Each prediction is compared with that example's actual target, the difference is squared, and the squared differences are averaged across the training examples. This average is the empirical squared loss for that single-feature assessment.
Squaring makes the score reflect the size of prediction errors without allowing positive and negative differences to cancel each other. In this loss-based scoring convention, a smaller average squared loss means the one-feature predictor made smaller errors on the assessed training examples, so the feature receives the more favorable loss score.
Comparing Two Loss Scores
A filter method evaluates two features with the same error-based quality measure. Feature A receives an empirical squared loss of 4, while Feature B receives an empirical squared loss of 9.
Identify what each number represents: Each number is the average of squared prediction errors produced by a predictor trained using only the corresponding feature.
Compare the losses: The loss for Feature A is smaller than the loss for Feature B.
Interpret the comparison: Under this loss-based scoring convention, the predictor using Feature A made smaller average squared errors on the assessed examples.
Use the result in selection: If the filter retains features with lower empirical squared loss, Feature A is the stronger candidate in this comparison.
Feature A receives the more favorable score under a lower-is-better empirical squared-loss rule.
From Scores to Selected Features
After every candidate feature has been assessed independently, the filter method has one score per feature. Selection happens at this later stage. The scores are compared, often by ranking them or applying a rule for retaining candidates, and the strongest candidates are kept according to the chosen quality measure.
What do you think happens?
A filter method has already computed one score for every feature. Does it need to retrain one combined model before it can compare the features?
Reveal answer
Answer: No, the independently produced scores can be compared afterward.
The filter procedure separates assessment from selection. Assessment produces one result per feature, and selection happens afterward when those results are compared.
Mistakes in Feature Scoring
Treating a filter method as though it evaluates only the full feature set.
The defining procedure evaluates each feature independently before selection.
Fix:
Assess one feature at a time and record one quality result for each feature.Confusing a feature's score with the final selection decision.
Assessment produces scores first. Selection occurs afterward when the scores are compared.
Fix:
Describe the score as an assessment result and apply the selection rule only after all relevant scores have been produced.Training the scoring predictor from more than the feature under assessment.
The predictive quality measure is based on a predictor trained solely from the feature being assessed.
Fix:
Use only the selected feature as the predictor's input during that feature's assessment.Reading squared loss as an unprocessed prediction difference.
Empirical squared loss is formed from squared prediction errors averaged across the training examples.
Fix:
Compare each prediction with its matching target, square the difference, and average the squared differences.
When explaining a filter result, state both the quality measure and the selection direction. For example, say that features were compared using empirical squared loss and that lower loss was treated as more favorable. This makes clear how the numerical scores led to the retained set.
Practice the Assessment Sequence
A filter method is assessing Feature C with an error-based predictive quality measure. Explain the sequence from the feature values and matching target values to the final feature score, and then explain what happens before Feature C can be selected.
Hints
- Start with the values of Feature C across the training examples.
- Mention that the predictor is trained using only Feature C.
- Describe how predictions are compared with matching targets.
- For squared loss, include squaring the errors and averaging them.
- Selection occurs only after Feature C's score is compared with the scores of other features.
- A filter method reduces a feature set by scoring features independently. A predictive quality measure can train a predictor using only the feature under assessment. In the regression illustration, empirical squared loss compares each prediction with its matching target, squares the error, and averages across the training examples. The resulting score is not the selection decision itself; scores are compared afterward to determine which features are retained.
Key Takeaways
- Filter methods assess each candidate feature independently before features are combined in a larger predictive model.
- A one-feature predictor can provide an error-based quality measure for the feature used to train it.
- Empirical squared loss is formed by squaring prediction errors and averaging them across the assessed training examples.
- Feature assessment produces scores; selection happens afterward when those scores are compared.
- The quality measure is not universal, so the scoring rule and its preferred direction should be stated clearly.