Concepts / Training and Evaluating Machine Learning Models

Training and Evaluating Machine Learning Models

A machine learning project needs an observable definition of success before progress can be measured.

  • Programming

Start with a Target

A machine learning project cannot improve toward an undefined target. Before choosing a model or deciding how to train it, define what success means and how success will be observed. That observable definition is the measure of success: a metric used to evaluate model performance.

The metric should connect to the higher-level goal, such as the success of a business, rather than being selected only because it is convenient to calculate.

What do you think happens?

A team has selected a model but has not defined what counts as success. Can it yet determine whether one training approach is better than another?

  • Yes, because any available metric is sufficient
  • No, because progress needs an observable definition of success
  • Yes, because the model architecture defines success
Reveal answer

Answer: No, because progress needs an observable definition of success

Without a defined target and a way to observe it, the project has no grounded basis for judging improvement.

Connect Goals to Metrics

Choosing a measure of success has two parts. First, identify the higher-level goal. Second, select a metric that provides an observable connection to that goal. The metric supplies the notion of performance that matters and can guide the choice of loss function.

From a project goal to an evaluation target

A machine learning team wants to make progress on a project whose success is defined at a higher level. It must decide what to measure before comparing training approaches.

Identify the goal: State the higher-level outcome that the project is intended to support.

Define observable success: Translate that goal into a measure that can be observed when evaluating model performance.

Use the measure consistently: Use the resulting metric to compare progress and guide the choice of loss function.

The project now has an observable target instead of an undefined idea of improvement.

Separate Metrics from Protocols

Metric choice and evaluation-protocol choice are related, but they answer different questions. The metric asks, What notion of performance matters? The evaluation protocol asks, What procedure will we use to assess that performance?

DecisionQuestion it answersWhat primarily guides it
Measure of successWhat counts as good model performance?Alignment with the higher-level goal and problem type
Evaluation protocolHow will current performance be assessed?Data availability

Metric and protocol decisions support each other but should not be confused.

definesevaluated throughprovides data forproduces a model forproducesHigher-level goalSuccess metricData procedurehold-out or foldsTrainingValidationPerformance estimate
How does an evaluation protocol move from defining a target to estimating model performance?

Match Metrics to Problem Types

Metric choice depends on the problem type. Balanced-classification, class-imbalanced, ranking, and multilabel problems are different contexts for deciding what success should mean. The correct choice is therefore not simply the metric that is easiest to calculate; it is the metric that represents the relevant notion of performance for the problem.

Problem contextMetric-selection question
Balanced classificationWhich measure represents success for this classification task?
Class-imbalanced classificationWhich measure represents success when the class distribution affects what should count as good performance?
RankingWhich measure represents success in the ordering produced by the model?
MultilabelWhich measure represents success when an example can involve multiple labels?

Before naming a metric, describe the task and the higher-level outcome it supports. Then verify that the metric measures that outcome rather than merely offering a convenient numerical score.

Hold Out a Validation Set

Hold-out validation sets aside part of the available data for validation. The remaining data is used for training, while the reserved portion provides the validation assessment. This is the simple choice when plenty of data is available, because reserving some examples is less likely to make the remaining training data too small for the task.

A single hold-out split

A project has plentiful data and needs a straightforward way to assess a model during development.

Reserve data: Set aside part of the available data as the validation set.

Train: Use the remaining data for training.

Validate: Assess the trained model using the reserved validation portion.

The project obtains a validation assessment while retaining a large training portion.

Rotate Through K Folds

K-fold cross-validation is recommended when there are too few samples for hold-out validation to be reliable. The data is divided into folds. During the procedure, each fold can take a turn as the evaluation portion, while the other folds provide the remaining data used in that evaluation cycle.

remaining dataremaining dataremaining datatrain then evaluateFold 1validationOther foldstraining dataEvaluation scorefrom each cycleFold 2validationFold 3validation
How does each fold take a turn serving as validation data while the other folds are used for training?

The important state change in K-fold cross-validation is the role of each fold. A fold that is used for evaluation in one cycle is part of the remaining data in other cycles. This rotation allows the procedure to evaluate the model on each fold rather than relying on one fixed hold-out portion.

Repeat the Fold Assignments

Iterated K-fold validation repeats K-fold cross-validation multiple times. Its defining change is repetition: instead of carrying out K-fold cross-validation only once, the evaluation procedure is carried out again to obtain results from multiple repetitions. The source recommends this approach for highly accurate model evaluation when little data is available.

producesproducesproducescombined for evaluationK-fold run 1fold assignmentsMultiple resultsfrom repeated proceduresEvaluation estimateK-fold run 2fold assignmentsK-fold run 3fold assignments
What changes when K-fold cross-validation is repeated, and how does repetition produce results from multiple evaluation runs?
ProtocolData situationDefining operation
Hold-out validationPlentiful dataSet aside one validation portion
K-fold cross-validationToo few samples for reliable hold-out validationRotate the evaluation role across folds
Iterated K-fold validationLittle data and a need for highly accurate evaluationRepeat K-fold cross-validation

Choose the Procedure

MEDIUM

A project has limited data and the team believes that one hold-out split may not provide a reliable assessment. Which evaluation protocol should it consider first, and what change would justify moving to the iterated version?

Hints
  • Use data availability as the main decision factor.
  • Distinguish one pass through the folds from repeating that procedure.
  • The iterated version is associated with highly accurate evaluation when little data is available.

Selecting an evaluation protocol

Compare three project situations: plentiful data, too few samples for reliable hold-out validation, and little data with a need for highly accurate evaluation.

Plentiful data: Choose hold-out validation because reserving a portion is less likely to make the training data too small.

Too few samples: Choose K-fold cross-validation so each fold can take a turn as the evaluation portion.

Little data with high accuracy required: Choose iterated K-fold validation because it repeats K-fold cross-validation.

Protocol selection follows the available data and the required accuracy of the evaluation, while metric selection remains a separate decision tied to the problem and higher-level goal.

Mistakes to Avoid

  • Choosing a metric only because it is convenient to calculate.

    The project may measure the wrong target carefully.

    Fix: Define observable success in relation to the higher-level goal before selecting the metric.

  • Treating the metric and the evaluation protocol as the same decision.

    A strong metric without a suitable evaluation protocol leaves the progress estimate poorly supported.

    Fix: Choose the metric for what performance means and the protocol for how performance will be assessed.

  • Using hold-out validation automatically when data is limited.

    The source identifies K-fold cross-validation as the recommended choice when hold-out validation is not reliable because of limited samples.

    Fix: Consider K-fold cross-validation, or iterated K-fold validation when highly accurate evaluation is needed with little data.

  • Describing iterated K-fold validation as merely another name for K-fold cross-validation.

    The defining feature of the iterated approach is that K-fold cross-validation is repeated.

    Fix: Reserve the term iterated K-fold validation for repeated K-fold evaluation procedures.

Key Takeaways

  1. Define an observable measure of success before choosing a model or deciding how to train it.
  2. Align the metric with the higher-level goal and the problem type; metric choice and evaluation-protocol choice are different decisions.
  3. Use hold-out validation when data is plentiful.
  4. Use K-fold cross-validation when there are too few samples for hold-out validation to be reliable.
  5. Use iterated K-fold validation when little data is available and highly accurate evaluation is needed.

Key Takeaways

  • A machine learning project needs an observable definition of success before progress can be measured.
  • A metric defines the notion of performance, while an evaluation protocol defines the procedure used to assess it.
  • Metric choice depends on the problem type and alignment with the higher-level goal.
  • Evaluation-protocol choice depends largely on data availability: hold-out for plentiful data, K-fold for limited data, and iterated K-fold for highly accurate evaluation with little data.