Feature Engineering
A machine learning problem is defined by its inputs, predicted outputs, and problem type.
From Raw Data to a Learnable Task
A machine learning project does not begin with choosing an algorithm. It begins by deciding what information is available, what the model should predict, and what kind of prediction problem is being addressed. Feature engineering belongs inside this broader problem-definition process: it transforms raw data into more meaningful features before the data enters a model.
The problem type matters because it guides choices such as model architecture and loss function. Possible types include binary classification, multiclass classification, scalar regression, vector regression, multiclass multilabel classification, clustering, generation, and reinforcement learning. The same raw data can support different tasks depending on what output is being predicted.
Testing the Two Working Hypotheses
Having examples of inputs X and targets Y does not prove that a machine learning problem is solvable. Before choosing a model, two working hypotheses must be examined. First, the inputs must contain useful information about the outputs. Second, the available data must represent the relationship that you want the model to learn.
Checking Whether a Prediction Task Is Well Defined
Suppose a team wants to predict an output from a collection of available inputs.
Identify the inputs: List the information that will be available when the model must make a prediction.
Identify the output: State exactly what the model is expected to predict.
Check information content: Ask whether the inputs contain useful information about the output. Merely pairing inputs with targets does not establish this.
Check representation: Ask whether the available examples represent the relationship that the model is intended to learn.
Choose the problem type: Classify the task in a way that can guide model architecture and loss-function choices.
A dataset becomes a meaningful basis for a machine learning problem only after the inputs, predicted outputs, problem type, information content, and relationship representation have been examined.
When the Past Stops Representing the Future
Machine learning relies on patterns in past training data to support predictions about future behavior. This makes a further assumption important: the future must resemble the relevant patterns represented in the past data. When the relationship between inputs and outputs changes over time, the problem can become nonstationary.
Nonstationarity is not simply a property of having data from different dates. The important issue is whether the relationship relevant to prediction changes over time. Therefore, the relevant time scale matters: a relationship may appear stable over one period but change over another. A model trained on historical patterns is limited when those patterns no longer describe the future behavior it must predict.
What do you think happens?
If the inputs and targets are correctly recorded but the relationship between them changes after training, should the original training pattern automatically be expected to remain reliable?
Reveal answer
Answer: No, because the future relationship may no longer match the historical pattern.
Machine learning uses patterns in past training data to support predictions about future behavior. A changing input-output relationship can make the problem nonstationary and weaken that assumption.
Transforming Data Before Modeling
Feature engineering is the use of non-learned transformations to turn raw data into more meaningful features before the data enters a machine learning model.
A model does not always receive data in the form that makes its task easiest. Feature engineering improves the input by using knowledge about the data and the machine learning algorithm. The goal is not merely to rearrange data; it is to express the task in a simpler and more useful way.
Making a Task Simpler Through Representation
Consider a raw dataset in which the information needed for a prediction is present but not expressed in the most useful form.
Begin with the raw representation: The model receives the original data, even though the relevant pattern may be difficult to express in that form.
Use problem knowledge: Study the task deeply enough to recognize which parts of the raw data are meaningful for the prediction.
Apply a non-learned transformation: Create a feature representation that expresses the relevant information more directly before the data enters the model.
Reconsider the learning problem: The transformed representation can change a difficult machine learning problem into a much simpler one.
Feature engineering simplifies learning by changing how the information is represented, not by changing the prediction goal.
Good feature engineering requires understanding both the problem and the data. You must recognize which parts of the raw data matter for the task and how to represent them so that the model can use them more effectively.
Shallow Algorithms and Deep Models
| Context | Role of feature engineering | Reason |
|---|---|---|
| Classical shallow algorithms | Critical preparation step | These algorithms did not have hypothesis spaces rich enough to learn useful features by themselves. |
| Deep learning | Reduced need, but still valuable | Deep models can automatically extract useful features from raw data, while manually designed features can still produce elegant solutions, use fewer resources, or solve a problem with less data. |
Before deep learning, how data was presented to a classical shallow algorithm was essential to success because the algorithm could not learn useful feature representations by itself. Deep learning changed this balance: deep models can automatically extract useful features from raw data, which removes the need for most manual feature engineering.
Automatic feature extraction does not make manually designed features irrelevant. A good feature can still produce a more elegant solution, use fewer resources, or solve a problem with far less data. In both settings, understanding the problem deeply remains useful because that understanding can reveal a simpler representation.
Mistakes in Problem Definition
Assuming that input-target pairs automatically make a problem solvable.
The inputs may not contain useful information about the outputs, or the available data may not represent the relationship that the team wants to learn.
Fix:
Test both working hypotheses before selecting a model.Ignoring the problem type.
Problem type guides choices such as model architecture and loss function.
Fix:
State the predicted output and classify the structure of the task.Treating historical patterns as permanently reliable.
Changing relationships can make a problem nonstationary, so past data may no longer represent future behavior.
Fix:
Consider changing relationships and the relevant time scales.Assuming deep learning eliminates the value of feature engineering.
Deep learning reduces the need for manual features but does not make them irrelevant.
Fix:
Consider whether a designed feature could produce a simpler solution, use fewer resources, or require less data.
Feature Engineering Practice
Describe how you would examine a proposed machine learning task before choosing a model. Include the available inputs, the predicted outputs, the problem type, the two working hypotheses, the possibility of a changing relationship over time, and one reason a feature transformation might help.
Hints
- Begin by separating what is available from what must be predicted.
- Ask whether the inputs contain useful information about the outputs.
- Ask whether the available examples represent the relationship you want to learn.
- Consider whether the relationship could change over time.
- Explain how a non-learned transformation could express the relevant information more meaningfully.
- A machine learning problem is defined by its inputs, predicted outputs, and problem type. Its feasibility depends on data availability, useful information in the inputs, and whether the available examples represent the relationship to be learned. Because machine learning uses past patterns to support future predictions, changing relationships can create nonstationary problems. Feature engineering applies non-learned transformations before modeling so raw data becomes a more meaningful representation. Classical shallow algorithms depended heavily on this preparation, while deep learning can learn many features automatically. Even so, designed features can still simplify a task, reduce resources, or reduce the amount of data needed.
Key Takeaways
- Define a machine learning task by its inputs, predicted outputs, and problem type.
- Check that the inputs contain useful information and that the available data represents the relationship to be learned.
- Remember that changing input-output relationships over time can make a problem nonstationary.
- Use feature engineering to transform raw data into more meaningful features before modeling.
- Deep learning reduces manual feature engineering but does not eliminate the value of well-designed features.