Evaluating Machine Learning Models
A machine learning problem is defined by its inputs, predicted outputs, and problem type.
From Task to Prediction
Before evaluating a machine learning model, first make the task precise. A machine learning problem is defined by three connected elements: the inputs provided to the model, the outputs it is expected to predict, and the problem type. This definition determines what the model is supposed to learn and provides the basis for judging whether its predictions are useful.
A Task Definition in Practice
Classifying Support Requests
Suppose a team wants a machine learning system to assign each support request to one of several categories.
Identify the inputs: The request information supplied to the model is the input. The task definition must state what information the model can use.
Identify the predicted output: The category assigned to the request is the predicted output.
Identify the problem type: Because the output is one category selected from several categories, the task is a multiclass classification problem.
Plan evaluation: The model's predicted category can be compared with known target categories for examples where those targets are available.
The task definition contains inputs, a categorical predicted output, and a multiclass classification problem type.
The problem type is not merely a label. It guides choices such as the model architecture and loss function. Possible problem types include binary classification, multiclass classification, scalar regression, vector regression, multiclass multilabel classification, clustering, generation, and reinforcement learning.
| Task element | Question to ask |
|---|---|
| Inputs | What information is supplied to the model? |
| Predicted outputs | What should the model produce? |
| Problem type | What kind of prediction or learning task is this? |
A compact checklist for defining a machine learning problem.
The Two Working Hypotheses
Defining a machine learning problem involves two important hypotheses. First, the selected inputs contain useful information about the outputs. If the inputs contain no useful signal about the targets, collecting input-target pairs does not make reliable prediction possible. Second, the available data represents the relationship that the model needs to learn. Data from the past is useful only under the assumption that future cases will resemble the patterns represented in that data.
Data as the First Limit
Data availability is often the first practical limitation on a machine learning problem. If relevant examples are unavailable, the problem may not be possible to attempt in the intended way. However, having examples of inputs and targets still does not guarantee success. The examples must provide information about the relationship the model is expected to learn.
Check data availability before choosing a model. Ask what examples exist, whether the examples contain the relevant inputs and targets, and whether they represent the relationship that matters for the intended task.
When Relationships Move
A problem can become nonstationary when the relationship between inputs and outputs changes over time. A model may have learned a relationship that was useful for one period, while later data follows a different relationship. In that situation, evaluating the model requires attention to the changing relationship and to the relevant time scales.
Imagine that the relationship between a set of inputs and its target is stable during an earlier period but changes later. A model trained on the earlier relationship may no longer be evaluated against the same underlying pattern. The important question is not only how the model behaved in the earlier data, but also whether the relationship remains relevant at the time and scale where the model is used.
Checking a Model's Evidence
Evaluation should be connected to the task definition. Start with the inputs the model receives and the outputs it is expected to predict. Then compare its predictions with known outcomes when those outcomes are available. The comparison is meaningful only if the inputs contain useful information about the outputs and the data used for evaluation represents the relationship that the model is meant to learn.
- State the inputs available to the model.
- State the predicted outputs.
- Identify the problem type.
- Check whether relevant input-target data is available.
- Ask whether the inputs contain useful information about the outputs.
- Ask whether the available data represents the relationship to be learned.
- Check whether the relationship may change over time.
Assuming that input-target pairs automatically make a problem solvable.
Examples alone do not prove that the inputs support prediction.
Fix:
Check whether the inputs contain useful information about the outputs before choosing a model.Ignoring the problem type.
The problem type guides choices such as model architecture and loss function.
Fix:
Classify the task before selecting a modeling approach.Assuming that an old relationship will always remain valid.
A changing relationship can make a problem nonstationary.
Fix:
Consider changing relationships and the relevant time scales.
Practice the Definition
For a proposed machine learning task, write down its inputs, predicted outputs, and problem type. Then answer two questions: Do the inputs contain useful information about the outputs? Does the available data represent the relationship the model needs to learn? Finally, decide whether a relationship that changes over time could affect evaluation.
Hints
- Do not begin with the model architecture. Begin with the task definition.
- Separate the existence of examples from the question of whether those examples contain useful predictive information.
- Consider whether the relationship between inputs and outputs stays relevant over the time period of interest.
What do you think happens?
A dataset contains many input-target examples. Is that alone enough to conclude that the machine learning problem is solvable?
Reveal answer
Answer: No, the inputs must contain useful information about the outputs and the data must represent the relationship to be learned.
The source material identifies data availability as a practical limitation but also warns that input-target examples alone do not prove that the inputs contain enough information for prediction.
Key Takeaways
- A machine learning problem is defined by its inputs, predicted outputs, and problem type.
- Data availability is often the first practical limitation, but examples alone do not guarantee that prediction is possible.
- The two central hypotheses are that the inputs contain useful information about the outputs and that the available data represents the relationship to be learned.
- A problem is nonstationary when the relationship between inputs and outputs changes over time.
- Machine learning depends on future behavior resembling patterns found in past training data.
Key Takeaways
- Define every machine learning task through its inputs, predicted outputs, and problem type.
- Treat data availability and data relevance as separate checks.
- Remember the two hypotheses: inputs contain useful information about outputs, and available data represents the relationship to be learned.
- Watch for nonstationarity when relationships change over time.
- Evaluate a model with awareness that its usefulness depends on future cases resembling patterns in past data.