Data Preprocessing
Feature engineering transforms raw data into more meaningful features before modeling.
Why Representation Matters
A machine learning model does not always receive data in the form that makes its task easiest. Feature engineering addresses this problem by transforming raw data into more meaningful features before the data enters the model. The central idea is simple: instead of asking the model to work directly with every raw detail, prepare an input representation that expresses the information in a form more useful for the task.
Feature engineering is a preprocessing step. Its transformations are designed before the model receives the data; they are not features learned by the model during training.
From Raw Data to Features
The transformation stage uses knowledge about the data and the machine learning algorithm to create a more meaningful representation. The transformed features then enter the model. This changes what the model has to work with, rather than changing the model itself.
A Simpler Representation
Consider a generated example in which raw data contains many separate observations about an item. A preprocessing designer might transform those observations into a smaller set of features that captures the information most relevant to the prediction task. The exact transformation depends on the problem, but the reasoning follows the same pattern: identify meaningful information, represent it explicitly, and pass that representation to the model.
Choosing a More Meaningful Representation
A model receives raw data that does not express the prediction task in an easy-to-use form. How should feature engineering be used?
Identify the difficulty: Examine what makes the raw representation difficult for the model and determine which parts of the data are meaningful for the task.
Design a transformation: Use knowledge about the data and the algorithm to create features that express the meaningful information more directly.
Place the transformation before the model: Apply the non-learned transformation before the data enters the model, so the model receives the engineered features rather than only the original representation.
Evaluate the new representation: The goal is not transformation for its own sake. The goal is to express the problem more simply.
Feature engineering can turn a difficult machine learning problem into a simpler one by changing how the relevant information is represented before modeling.
The difficult part is often not performing a transformation but choosing one that reflects the problem. Useful feature design depends on understanding which aspects of the raw data are meaningful for the task.
Shallow Models and Deep Models
Before deep learning, feature engineering was critical. Classical shallow algorithms did not have hypothesis spaces rich enough to learn useful features by themselves, so the way data was presented to the algorithm was essential to success. Human-designed features supplied the representation that the shallow algorithm needed.
Deep learning changes the balance. Deep learning models can automatically extract useful features from raw data, which removes the need for most manual feature engineering. However, automatic extraction does not make human-designed transformations worthless. A good manually designed feature can still produce a more elegant solution, use fewer resources, or solve a problem with far less data.
| Aspect | Classical shallow algorithms | Deep learning |
|---|---|---|
| Feature extraction | Useful features generally had to be designed before modeling. | Useful features can be extracted automatically from raw data. |
| Role of representation | How the data was presented was essential to algorithm success. | The model can learn useful internal representations. |
| Value of manual features | Historically critical. | Reduced in many cases, but not eliminated. |
| Possible benefit of good manual features | Provide the representation the algorithm needs. | Can produce a more elegant solution, use fewer resources, or require less data. |
Designing Useful Features
Begin with the problem rather than with a favorite transformation. Ask what information is meaningful for the task and whether the raw representation expresses that information clearly. Then use knowledge about both the data and the machine learning algorithm to design a representation that makes the task easier.
Common Mistakes
Assuming feature engineering is unnecessary whenever deep learning is used.
Deep learning reduces the need for most manual feature engineering, but good features can still provide a more elegant solution, use fewer resources, or solve a problem with less data.
Fix:
Consider whether problem knowledge can produce a useful representation before deciding that automatic extraction is sufficient.Creating a transformation without understanding the prediction problem.
Useful feature design usually requires a deep understanding of the problem and of how meaningful information should be represented.
Fix:
Start by identifying the parts of the raw data that are relevant to the task.Thinking that feature engineering changes the model's internal learning process.
Feature engineering creates more meaningful features before the data enters the model; it is distinct from features extracted automatically by a model.
Fix:
Separate the non-learned preprocessing stage from the model's own learned representations.Treating raw data as automatically suitable for a model.
A model does not always receive data in the most useful form, and a good transformation can simplify a difficult machine learning problem.
Fix:
Ask whether a different representation would express the relevant information more directly.
Practice
A team has raw data that makes its prediction task difficult. The team is considering a manually designed transformation before the model receives the data. Explain why this could help, and explain why the same idea can still be valuable when the team uses deep learning.
Hints
- Describe what happens before the data enters the model.
- Explain how a more meaningful representation can simplify the task.
- Contrast automatic feature extraction in deep learning with the possible benefits of human-designed features.
What do you think happens?
A deep learning model can automatically extract useful features. Does that mean manually designed features have no remaining value?
Reveal answer
Answer: No, good manual features can still produce a more elegant solution, use fewer resources, or require less data.
Deep learning reduces the need for most manual feature engineering, but it does not eliminate the value of problem-specific representations.
Key Takeaways
- Feature engineering transforms raw data into more meaningful features before modeling.
- The goal is to express the problem in a simpler form for the machine learning model.
- Useful feature design requires understanding which parts of the raw data matter and how to represent them.
- Feature engineering was historically critical for classical shallow algorithms because they could not learn useful features by themselves.
- Deep learning can automatically extract features, but manually designed features can still reduce resources, require less data, or provide a more elegant solution.
Key Takeaways
- Feature engineering is a preprocessing stage that creates more meaningful features before data enters a model.
- A well-designed representation can turn a difficult machine learning problem into a simpler one.
- Classical shallow algorithms depended heavily on human-designed features because they could not learn useful features by themselves.
- Deep learning reduces the need for manual feature engineering by extracting features automatically.
- Manual feature engineering remains useful when it produces a more elegant solution, uses fewer resources, or works with less data.