Concepts / Evaluating Machine Learning Models

Evaluating Machine Learning Models

A machine learning problem is defined by its inputs, predicted outputs, and problem type.

  • Programming

From Task to Prediction

Before evaluating a machine learning model, first make the task precise. A machine learning problem is defined by three connected elements: the inputs provided to the model, the outputs it is expected to predict, and the problem type. This definition determines what the model is supposed to learn and provides the basis for judging whether its predictions are useful.

used to predictdefines the prediction taskguides model choicesInputsinformation provided to themodelProblem typeclassification, regression,or another typePredicted outputstargets the model shouldproduce
How are the inputs, predicted outputs, and problem type connected when defining a machine learning task?

A Task Definition in Practice

Classifying Support Requests

Suppose a team wants a machine learning system to assign each support request to one of several categories.

Identify the inputs: The request information supplied to the model is the input. The task definition must state what information the model can use.

Identify the predicted output: The category assigned to the request is the predicted output.

Identify the problem type: Because the output is one category selected from several categories, the task is a multiclass classification problem.

Plan evaluation: The model's predicted category can be compared with known target categories for examples where those targets are available.

The task definition contains inputs, a categorical predicted output, and a multiclass classification problem type.

The problem type is not merely a label. It guides choices such as the model architecture and loss function. Possible problem types include binary classification, multiclass classification, scalar regression, vector regression, multiclass multilabel classification, clustering, generation, and reinforcement learning.

Task elementQuestion to ask
InputsWhat information is supplied to the model?
Predicted outputsWhat should the model produce?
Problem typeWhat kind of prediction or learning task is this?

A compact checklist for defining a machine learning problem.

The Two Working Hypotheses

Defining a machine learning problem involves two important hypotheses. First, the selected inputs contain useful information about the outputs. If the inputs contain no useful signal about the targets, collecting input-target pairs does not make reliable prediction possible. Second, the available data represents the relationship that the model needs to learn. Data from the past is useful only under the assumption that future cases will resemble the patterns represented in that data.

contains information aboutis related tosupportssupportsInputsUseful relationshipinputs contain informationabout outputsReliable predictionOutputsFuture dataresembles representedpatterns
What are the two assumptions about the relationship between inputs, outputs, and future data when defining a machine learning problem?

Data as the First Limit

Data availability is often the first practical limitation on a machine learning problem. If relevant examples are unavailable, the problem may not be possible to attempt in the intended way. However, having examples of inputs and targets still does not guarantee success. The examples must provide information about the relationship the model is expected to learn.

requiressupportsmay facecreatesLearning taskRelevant dataavailable examplesAttemptable problemMissing relevant dataPractical limitation
How does the availability or absence of relevant data determine which machine learning problems can be attempted?

Check data availability before choosing a model. Ask what examples exist, whether the examples contain the relevant inputs and targets, and whether they represent the relationship that matters for the intended task.

When Relationships Move

A problem can become nonstationary when the relationship between inputs and outputs changes over time. A model may have learned a relationship that was useful for one period, while later data follows a different relationship. In that situation, evaluating the model requires attention to the changing relationship and to the relevant time scales.

relationshipdifferent relationshippast patternchanges evaluationInputsearly periodOutputsearly relationshipInputslater periodOutputschanged relationshipModel reliabilityrequires attention to time
How can a changing relationship between inputs and outputs over time make a machine learning model stop working reliably?

Imagine that the relationship between a set of inputs and its target is stable during an earlier period but changes later. A model trained on the earlier relationship may no longer be evaluated against the same underlying pattern. The important question is not only how the model behaved in the earlier data, but also whether the relationship remains relevant at the time and scale where the model is used.

Checking a Model's Evidence

Evaluation should be connected to the task definition. Start with the inputs the model receives and the outputs it is expected to predict. Then compare its predictions with known outcomes when those outcomes are available. The comparison is meaningful only if the inputs contain useful information about the outputs and the data used for evaluation represents the relationship that the model is meant to learn.

  1. State the inputs available to the model.
  2. State the predicted outputs.
  3. Identify the problem type.
  4. Check whether relevant input-target data is available.
  5. Ask whether the inputs contain useful information about the outputs.
  6. Ask whether the available data represents the relationship to be learned.
  7. Check whether the relationship may change over time.
  • Assuming that input-target pairs automatically make a problem solvable.

    Examples alone do not prove that the inputs support prediction.

    Fix: Check whether the inputs contain useful information about the outputs before choosing a model.

  • Ignoring the problem type.

    The problem type guides choices such as model architecture and loss function.

    Fix: Classify the task before selecting a modeling approach.

  • Assuming that an old relationship will always remain valid.

    A changing relationship can make a problem nonstationary.

    Fix: Consider changing relationships and the relevant time scales.

Practice the Definition

MEDIUM

For a proposed machine learning task, write down its inputs, predicted outputs, and problem type. Then answer two questions: Do the inputs contain useful information about the outputs? Does the available data represent the relationship the model needs to learn? Finally, decide whether a relationship that changes over time could affect evaluation.

Hints
  • Do not begin with the model architecture. Begin with the task definition.
  • Separate the existence of examples from the question of whether those examples contain useful predictive information.
  • Consider whether the relationship between inputs and outputs stays relevant over the time period of interest.

What do you think happens?

A dataset contains many input-target examples. Is that alone enough to conclude that the machine learning problem is solvable?

  • Yes, examples always prove solvability.
  • No, the inputs must contain useful information about the outputs and the data must represent the relationship to be learned.
  • Yes, if the dataset is large enough.
Reveal answer

Answer: No, the inputs must contain useful information about the outputs and the data must represent the relationship to be learned.

The source material identifies data availability as a practical limitation but also warns that input-target examples alone do not prove that the inputs contain enough information for prediction.

Key Takeaways

  1. A machine learning problem is defined by its inputs, predicted outputs, and problem type.
  2. Data availability is often the first practical limitation, but examples alone do not guarantee that prediction is possible.
  3. The two central hypotheses are that the inputs contain useful information about the outputs and that the available data represents the relationship to be learned.
  4. A problem is nonstationary when the relationship between inputs and outputs changes over time.
  5. Machine learning depends on future behavior resembling patterns found in past training data.

Key Takeaways

  • Define every machine learning task through its inputs, predicted outputs, and problem type.
  • Treat data availability and data relevance as separate checks.
  • Remember the two hypotheses: inputs contain useful information about outputs, and available data represents the relationship to be learned.
  • Watch for nonstationarity when relationships change over time.
  • Evaluate a model with awareness that its usefulness depends on future cases resembling patterns in past data.