Fighting Overfitting
A machine learning problem is defined by its inputs, predicted outputs, and problem type.
Before Choosing a Model
Overfitting is often discussed as a model problem, but a deeper problem can appear earlier: the machine learning task may not be defined in a way that makes reliable prediction possible. Before choosing a model, start by identifying what information is available, what the model should predict, and what kind of prediction task you are solving.
A machine learning problem is defined by its inputs, predicted outputs, and problem type.
Three Parts of a Task
A useful problem definition connects three parts. Inputs are the information supplied to the model. Predicted outputs are the results the model is expected to produce. The problem type describes the structure of that output and the task the model must perform. These parts are connected: the inputs must be considered in relation to the desired outputs, and the problem type guides choices such as model architecture and loss function.
Separating the Task Parts
Suppose a model receives available information and must produce a prediction. Define the task without choosing a model yet.
Identify the inputs: Write down the information that will be supplied to the model.
Identify the output: State exactly what the model is expected to predict.
Identify the problem type: Describe the structure of the prediction, such as a class, a scalar, a vector, or another listed task type.
Delay model selection: Only after the task is defined should the problem type guide choices such as model architecture and loss function.
A clear task definition states the inputs, the predicted outputs, and the problem type before a model is selected.
Data as a Limitation
Data availability is often the first practical limitation on a machine learning problem. To learn from examples, you need input examples and known target outputs. However, having input-target examples does not prove that the inputs contain enough information for prediction. It also does not prove that the available data represents the relationship you want to learn.
The Two Working Hypotheses
Defining a machine learning problem involves two important hypotheses. First, the inputs contain useful information about the outputs; otherwise, the model has no informative basis for making the prediction. Second, the available data represents the relationship that you want to learn, including the behavior that will matter when the model is used. These are hypotheses because examples alone do not automatically establish that either condition is true.
When Patterns Move
A problem can be nonstationary when the relationship between inputs and outputs changes over time. A model trained on earlier data may then make increasingly incorrect predictions because the pattern it learned no longer matches the current relationship. Nonstationary problems require attention to the changing relationship and to the relevant time scales.
When time matters, do not ask only whether the training examples are available. Ask whether the relationship between inputs and outputs is stable at the time scale relevant to use. If that relationship changes, treat the task as potentially nonstationary.
Why the Future Must Resemble the Past
Machine learning uses patterns found in past training data to make predictions about future behavior. This only works when the future retains enough resemblance to the patterns represented in the training data. If the future no longer resembles the past, the model's learned relationship may no longer provide useful predictions. This is the central limitation behind using historical examples to predict what comes next.
What do you think happens?
A model was trained on examples from an earlier period. If the relationship between its inputs and outputs changes later, what should you expect?
Reveal answer
Answer: The old model may make increasingly incorrect predictions.
The model learned a relationship from past data. When that relationship changes, the past pattern may no longer match current behavior.
Task Types and Model Choices
| Problem type | Predicted output structure | Why it matters |
|---|---|---|
| Binary classification | One of two classes | The task type helps guide model architecture and loss function. |
| Multiclass classification | One class among multiple classes | The task type helps guide model architecture and loss function. |
| Scalar regression | A scalar value | The task type helps guide model architecture and loss function. |
| Vector regression | A vector of values | The task type helps guide model architecture and loss function. |
| Multiclass multilabel classification | Multiple class labels | The task type helps guide model architecture and loss function. |
| Clustering | Groups discovered in data | The task type identifies a different machine learning task structure. |
| Generation | Generated output | The task type identifies a different machine learning task structure. |
| Reinforcement learning | A learning task based on interaction | The task type identifies a different machine learning task structure. |
Do not treat the problem type as a label added after modeling. It is part of the definition of the machine learning problem. Whether the task is binary classification, multiclass classification, scalar regression, vector regression, multiclass multilabel classification, clustering, generation, or reinforcement learning affects how the task is understood and guides choices such as model architecture and loss function.
Mistakes in Problem Definition
Assuming that input-target pairs automatically make the task solvable.
The inputs may not contain useful information about the outputs, and the available data may not represent the relationship that must be learned.
Fix:
Check both whether the inputs carry useful information and whether the data represents the relationship of interest.Ignoring the problem type.
Problem type guides choices such as model architecture and loss function.
Fix:
Name the problem type as part of the initial task definition.Treating historical patterns as permanently stable.
A changing relationship can make the problem nonstationary and cause increasingly incorrect predictions.
Fix:
Consider changing relationships and the relevant time scales before relying on past training data.
Apply the Definition
Describe a hypothetical machine learning task in four parts: the inputs, the predicted outputs, the problem type, and the assumption that must hold for past training data to help with future predictions. Then state one reason the task might fail even if input-target examples are available.
Hints
- Start by naming the information supplied to the model.
- Describe the output structure before naming the problem type.
- Check whether the inputs contain useful information about the outputs.
- Check whether the relationship represented by the data is likely to remain relevant in the future.
- A strong machine learning problem definition identifies the inputs, predicted outputs, and problem type. Data availability is only the starting point: the inputs must contain useful information, and the available data must represent the relationship to be learned. Machine learning relies on the future resembling patterns found in past training data. When the input-output relationship changes over time, the problem may be nonstationary and predictions may become increasingly incorrect.
Key Takeaways
- Define a machine learning problem by its inputs, predicted outputs, and problem type.
- Input-target examples do not guarantee that the inputs contain useful information or that the data represents the relationship to be learned.
- The two key hypotheses are that inputs contain useful information about outputs and that available data represents the relationship of interest.
- Future predictions rely on the future resembling patterns in past training data.
- Changing input-output relationships can create nonstationary problems and make predictions increasingly incorrect.