Training and Evaluating Machine Learning Models
A machine learning project needs an observable definition of success before progress can be measured.
Start with a Target
A machine learning project cannot improve toward an undefined target. Before choosing a model or deciding how to train it, define what success means and how success will be observed. That observable definition is the measure of success: a metric used to evaluate model performance.
The metric should connect to the higher-level goal, such as the success of a business, rather than being selected only because it is convenient to calculate.
What do you think happens?
A team has selected a model but has not defined what counts as success. Can it yet determine whether one training approach is better than another?
Reveal answer
Answer: No, because progress needs an observable definition of success
Without a defined target and a way to observe it, the project has no grounded basis for judging improvement.
Connect Goals to Metrics
Choosing a measure of success has two parts. First, identify the higher-level goal. Second, select a metric that provides an observable connection to that goal. The metric supplies the notion of performance that matters and can guide the choice of loss function.
From a project goal to an evaluation target
A machine learning team wants to make progress on a project whose success is defined at a higher level. It must decide what to measure before comparing training approaches.
Identify the goal: State the higher-level outcome that the project is intended to support.
Define observable success: Translate that goal into a measure that can be observed when evaluating model performance.
Use the measure consistently: Use the resulting metric to compare progress and guide the choice of loss function.
The project now has an observable target instead of an undefined idea of improvement.
Separate Metrics from Protocols
Metric choice and evaluation-protocol choice are related, but they answer different questions. The metric asks, What notion of performance matters? The evaluation protocol asks, What procedure will we use to assess that performance?
| Decision | Question it answers | What primarily guides it |
|---|---|---|
| Measure of success | What counts as good model performance? | Alignment with the higher-level goal and problem type |
| Evaluation protocol | How will current performance be assessed? | Data availability |
Metric and protocol decisions support each other but should not be confused.
Match Metrics to Problem Types
Metric choice depends on the problem type. Balanced-classification, class-imbalanced, ranking, and multilabel problems are different contexts for deciding what success should mean. The correct choice is therefore not simply the metric that is easiest to calculate; it is the metric that represents the relevant notion of performance for the problem.
| Problem context | Metric-selection question |
|---|---|
| Balanced classification | Which measure represents success for this classification task? |
| Class-imbalanced classification | Which measure represents success when the class distribution affects what should count as good performance? |
| Ranking | Which measure represents success in the ordering produced by the model? |
| Multilabel | Which measure represents success when an example can involve multiple labels? |
Before naming a metric, describe the task and the higher-level outcome it supports. Then verify that the metric measures that outcome rather than merely offering a convenient numerical score.
Hold Out a Validation Set
Hold-out validation sets aside part of the available data for validation. The remaining data is used for training, while the reserved portion provides the validation assessment. This is the simple choice when plenty of data is available, because reserving some examples is less likely to make the remaining training data too small for the task.
A single hold-out split
A project has plentiful data and needs a straightforward way to assess a model during development.
Reserve data: Set aside part of the available data as the validation set.
Train: Use the remaining data for training.
Validate: Assess the trained model using the reserved validation portion.
The project obtains a validation assessment while retaining a large training portion.
Rotate Through K Folds
K-fold cross-validation is recommended when there are too few samples for hold-out validation to be reliable. The data is divided into folds. During the procedure, each fold can take a turn as the evaluation portion, while the other folds provide the remaining data used in that evaluation cycle.
The important state change in K-fold cross-validation is the role of each fold. A fold that is used for evaluation in one cycle is part of the remaining data in other cycles. This rotation allows the procedure to evaluate the model on each fold rather than relying on one fixed hold-out portion.
Repeat the Fold Assignments
Iterated K-fold validation repeats K-fold cross-validation multiple times. Its defining change is repetition: instead of carrying out K-fold cross-validation only once, the evaluation procedure is carried out again to obtain results from multiple repetitions. The source recommends this approach for highly accurate model evaluation when little data is available.
| Protocol | Data situation | Defining operation |
|---|---|---|
| Hold-out validation | Plentiful data | Set aside one validation portion |
| K-fold cross-validation | Too few samples for reliable hold-out validation | Rotate the evaluation role across folds |
| Iterated K-fold validation | Little data and a need for highly accurate evaluation | Repeat K-fold cross-validation |
Choose the Procedure
A project has limited data and the team believes that one hold-out split may not provide a reliable assessment. Which evaluation protocol should it consider first, and what change would justify moving to the iterated version?
Hints
- Use data availability as the main decision factor.
- Distinguish one pass through the folds from repeating that procedure.
- The iterated version is associated with highly accurate evaluation when little data is available.
Selecting an evaluation protocol
Compare three project situations: plentiful data, too few samples for reliable hold-out validation, and little data with a need for highly accurate evaluation.
Plentiful data: Choose hold-out validation because reserving a portion is less likely to make the training data too small.
Too few samples: Choose K-fold cross-validation so each fold can take a turn as the evaluation portion.
Little data with high accuracy required: Choose iterated K-fold validation because it repeats K-fold cross-validation.
Protocol selection follows the available data and the required accuracy of the evaluation, while metric selection remains a separate decision tied to the problem and higher-level goal.
Mistakes to Avoid
Choosing a metric only because it is convenient to calculate.
The project may measure the wrong target carefully.
Fix:
Define observable success in relation to the higher-level goal before selecting the metric.Treating the metric and the evaluation protocol as the same decision.
A strong metric without a suitable evaluation protocol leaves the progress estimate poorly supported.
Fix:
Choose the metric for what performance means and the protocol for how performance will be assessed.Using hold-out validation automatically when data is limited.
The source identifies K-fold cross-validation as the recommended choice when hold-out validation is not reliable because of limited samples.
Fix:
Consider K-fold cross-validation, or iterated K-fold validation when highly accurate evaluation is needed with little data.Describing iterated K-fold validation as merely another name for K-fold cross-validation.
The defining feature of the iterated approach is that K-fold cross-validation is repeated.
Fix:
Reserve the term iterated K-fold validation for repeated K-fold evaluation procedures.
Key Takeaways
- Define an observable measure of success before choosing a model or deciding how to train it.
- Align the metric with the higher-level goal and the problem type; metric choice and evaluation-protocol choice are different decisions.
- Use hold-out validation when data is plentiful.
- Use K-fold cross-validation when there are too few samples for hold-out validation to be reliable.
- Use iterated K-fold validation when little data is available and highly accurate evaluation is needed.
Key Takeaways
- A machine learning project needs an observable definition of success before progress can be measured.
- A metric defines the notion of performance, while an evaluation protocol defines the procedure used to assess it.
- Metric choice depends on the problem type and alignment with the higher-level goal.
- Evaluation-protocol choice depends largely on data availability: hold-out for plentiful data, K-fold for limited data, and iterated K-fold for highly accurate evaluation with little data.