Understanding the Bias-Complexity Tradeoff in Machine Learning
Selecting features determines the dimensionality of the feature space.
The Feature Choice
Suppose you are designing a learning algorithm to predict whether a patient will suffer a heart attack. The available features are blood pressure, body-mass index, age, level of physical activity, and income. The design question is not only which learning algorithm to use. You must also decide which features the algorithm should use. Selecting features determines the dimensionality of the feature space, so feature selection is already a choice about model complexity.
Fewer selected features create a smaller feature space. More selected features create a larger feature space.
From Features to Dimensions
Each selected feature contributes a dimension to the feature space in this example. An algorithm using blood pressure and body-mass index has two selected features, so it works in a two-dimensional feature space. An algorithm using blood pressure, body-mass index, age, physical activity, and income has five selected features, so it works in a five-dimensional feature space. The difference between these choices is the complexity side of the bias-complexity tradeoff.
The Heart-Attack Comparison
Choosing Between Two and Five Features
Compare an algorithm that uses blood pressure and body-mass index with an algorithm that uses blood pressure, body-mass index, age, physical activity, and income.
Represent the first choice: The first algorithm selects two features, so its input lies in a two-dimensional feature space.
Represent the second choice: The second algorithm selects five features, so its input lies in a five-dimensional feature space.
Identify the tradeoff: The two-feature choice uses a smaller feature space and therefore a simpler representation. The five-feature choice uses a larger feature space and therefore allows a richer representation.
Avoid declaring an automatic winner: The correct analysis is to identify the advantages and disadvantages associated with each level of complexity rather than simply declaring one algorithm better.
The two-feature and five-feature algorithms differ in feature-space dimensionality. That difference creates a bias-complexity question: what may be lost by using fewer features, and what may become harder to estimate when using more?
| Choice | Selected features | Feature-space dimensionality | Possible advantage | Possible disadvantage |
|---|---|---|---|---|
| Smaller feature set | Blood pressure and body-mass index | 2D | A smaller hypothesis-search problem can reduce estimation error | The class may be unable to contain the optimal classifier, increasing approximation error |
| Larger feature set | Blood pressure, body-mass index, age, physical activity, and income | 5D | A richer class can reduce approximation error | Finite data may not be sufficient to choose reliably among many possibilities, increasing estimation error |
Hypothesis Classes
A hypothesis class is the set of possible classifiers from which the learning algorithm chooses.
Choosing a feature set affects the representation available to the learner, while choosing a hypothesis class determines which classifiers are available within that representation. A class that is too small may not contain a classifier close to the optimal one. A class that is very rich offers many possible classifiers, but the learner may have difficulty choosing reliably among them when the available sample is finite.
Two Sources of Error
Approximation error comes from a hypothesis class that is not rich enough to contain the optimal classifier. Estimation error is related to the finite sample size and the complexity of the hypothesis class. These errors respond differently when the class becomes richer.
Increasing the richness of a hypothesis class gives the learner more possible classifiers. This makes it less likely that the class itself will block a good solution, so approximation error decreases. However, the same flexibility creates more ways to fit the particular sample. Because estimation error depends on both finite sample size and hypothesis-class complexity, estimation error might increase. The source associates this increase with overfitting.
Moving toward a very small class creates the reverse pressure. The learner has fewer choices to distinguish among, which reduces estimation error. But the class may be unable to contain the optimal classifier, increasing approximation error. When the class is too restricted, the result may be underfitting.
Using Prior Knowledge
Theoretically, the ideal hypothesis class would contain only the Bayes optimal classifier. That would avoid choosing among unnecessary alternatives. In practice, this class is not generally available because the Bayes optimal classifier depends on the underlying distribution D, and that distribution is unknown. If the distribution were already known, there would be little need for learning.
Since the ideal class is unavailable, practical design uses knowledge about the problem to propose a reasonable class. The source uses a rectangle as an illustration of how domain knowledge can guide class design, not as a universally correct shape for every classification problem. The goal is to choose a class whose approximation error is not excessively high while its estimation error remains reasonable.
When choosing features or a hypothesis class, do not ask only whether the learner can represent more possibilities. Also ask whether the available sample is sufficient to choose reliably among those possibilities. Prior knowledge about the domain can help identify a class with an acceptable balance.
Common Reasoning Errors
Assuming that the five-feature algorithm is automatically better.
The larger feature space can reduce approximation error, but a richer class may increase estimation error when the sample is finite.
Fix:
Describe both pressures before deciding which choice is appropriate.Treating fewer features as automatically better.
A small class can have lower estimation error but higher approximation error if it cannot contain the optimal classifier.
Fix:
Recognize the possibility of underfitting when the class is too restricted.Confusing approximation error with estimation error.
Approximation error comes from a class that is not rich enough, while estimation error is related to finite sample size and hypothesis-class complexity.
Fix:
Identify whether the problem is the limits of the class or the reliability of selecting from that class using finite data.Assuming that a universally correct hypothesis-class shape is known.
The rectangle is an example of domain knowledge guiding class design, not a universally correct shape.
Fix:
Use prior knowledge to propose a reasonable class while recognizing that the optimal classifier is unknown.
Apply the Tradeoff
A heart-attack prediction system can use either blood pressure and body-mass index or all five available features. Explain how the two choices differ in feature-space dimensionality. Then state one possible advantage and one possible disadvantage of each choice using approximation error and estimation error.
Hints
- Count the selected features to identify the dimensionality.
- For the smaller choice, consider what happens when the hypothesis class is too restricted.
- For the larger choice, consider what happens when finite data must support a richer class.
What do you think happens?
What error pressure is most directly associated with making the hypothesis class richer?
Reveal answer
Answer: Approximation error may decrease, while estimation error might increase
A richer class offers more possible classifiers, making it less likely that the class excludes a good solution. But finite data may not be sufficient to choose reliably among the additional possibilities.
The Design Balance
- Selecting features determines the dimensionality of the feature space: two selected features create a two-dimensional space, while five create a five-dimensional space.
- A smaller feature set and hypothesis class can reduce estimation error but may increase approximation error if the class is too restricted.
- A richer feature set and hypothesis class can reduce approximation error but may increase estimation error because finite data may not support reliable selection among many possibilities.
- The ideal Bayes-optimal hypothesis class is generally unavailable because the underlying distribution is unknown.
- Prior knowledge about the problem can guide the design of a hypothesis class with a reasonable balance between approximation and estimation error.
Key Takeaways
- Feature selection changes the dimensionality of the learner's feature space.
- Using two features creates a simpler representation than using five features, but neither choice is automatically better.
- Richer hypothesis classes can lower approximation error while possibly raising estimation error.
- Smaller hypothesis classes can lower estimation error while possibly raising approximation error.
- Prior knowledge helps practitioners choose a hypothesis class when the optimal classifier and underlying distribution are unknown.