Classification and Ranking Problems
A machine learning project needs an observable definition of success before progress can be measured.
Defining Progress
A machine learning project cannot improve toward an undefined target. Before choosing a model or deciding how to train it, define what success means and how success will be observed. This observable definition becomes the measure of success: a metric used to evaluate model performance.
The useful question is not simply whether a model produces predictions. The useful question is whether those predictions support the higher-level goal of the project. A metric should connect directly to that goal, such as the success of a business, rather than being chosen only because it is convenient to calculate. The selected measure of success also guides the choice of loss function.
Classification and Ranking
The problem type affects the measure of success. The source distinguishes balanced-classification, class-imbalanced, ranking, and multilabel problems as different settings for metric choice. A classification problem concerns assigning items to classes, while a ranking problem concerns ordering items by their relative relevance or score. The important principle is that the metric must fit the prediction task and the structure of its labels.
| Problem setting | What metric choice must reflect |
|---|---|
| Balanced classification | The classification task |
| Class-imbalanced classification | The classification task and unequal class structure |
| Ranking | The ordering task |
| Multilabel | The multilabel structure |
The source identifies these settings as distinct metric-choice contexts but does not prescribe particular metric names.
Metric Versus Protocol
A metric produces the notion of performance that matters. An evaluation protocol establishes the procedure used to assess that performance.
These decisions are related but different. Choosing a strong metric without a suitable evaluation protocol leaves the progress estimate poorly supported. Choosing a protocol without a meaningful metric measures the wrong target carefully. A trustworthy evaluation therefore needs both a meaningful measure of success and a suitable procedure for estimating it.
Three Validation Strategies
Once the measure of success is defined, the next decision is how to measure current progress. The source identifies three common evaluation protocols: hold-out validation, K-fold cross-validation, and iterated K-fold validation. The main practical factor emphasized here is data availability.
| Protocol | How the available data is used | When the source recommends it |
|---|---|---|
| Hold-out validation | Part of the data is set aside as a validation set; the remaining data is used for training. | When plenty of data is available. |
| K-fold cross-validation | The data is divided into folds. Each fold can serve as the evaluation portion while the other folds provide the remaining data used in that evaluation cycle. | When there are too few samples for hold-out validation to be reliable. |
| Iterated K-fold validation | K-fold cross-validation is repeated multiple times. | When highly accurate model evaluation is needed and little data is available. |
Following the Folds
K-fold cross-validation divides the data into folds and evaluates the model on each fold. During an evaluation cycle, one fold serves as the evaluation portion while the other folds provide the remaining data used in that cycle. The process moves the evaluation role across the folds, allowing each fold to serve as the evaluation portion.
The key idea is reuse across evaluation cycles. A data point can belong to the evaluation portion in one cycle and to the remaining data used in another cycle. This makes K-fold cross-validation useful when there are too few samples for a single hold-out validation set to provide a reliable evaluation.
Worked Protocol Choice
Selecting a Protocol from Data Availability
A team has a machine learning problem and must choose an evaluation protocol. Consider three situations: plentiful data, too few samples for a reliable hold-out validation, and little data where highly accurate evaluation is required.
Situation 1: Choose hold-out validation when plenty of data is available. Setting aside part of the available data is less likely to make the remaining training data too small for the task.
Situation 2: Choose K-fold cross-validation when there are too few samples for hold-out validation to be reliable. The folds take turns serving as the evaluation portion.
Situation 3: Choose iterated K-fold validation when little data is available and highly accurate model evaluation is needed. This repeats K-fold cross-validation multiple times.
Final check: These choices concern the evaluation protocol, not the metric. The metric still must be selected to match the higher-level goal and the problem type.
Plentiful data points toward hold-out validation, limited data points toward K-fold cross-validation, and little data with a need for highly accurate evaluation points toward iterated K-fold validation.
Common Selection Mistakes
Choosing a metric only because it is convenient to calculate.
The model may appear to improve while progress is being measured against the wrong target.
Fix:
Define observable success in relation to the higher-level goal before choosing the metric.Confusing the metric with the evaluation protocol.
K-fold cross-validation describes how performance is assessed, not which notion of performance matters.
Fix:
Choose the metric for the goal and the protocol for the available data.Using one metric choice for every problem type.
The source identifies these as different contexts for metric choice.
Fix:
Match the measure of success to the prediction task and label structure.Using hold-out validation automatically when data is limited.
The source recommends K-fold cross-validation when hold-out validation is unreliable because of limited samples.
Fix:
Use K-fold cross-validation for too few samples, and consider iterated K-fold validation when little data requires highly accurate evaluation.
Decision Practice
A project team says, “We will use the easiest metric to calculate, and we will use hold-out validation because it is simple.” Identify the two decisions that must be reconsidered. Then state what information should guide each decision.
Hints
- Separate the question of what counts as success from the question of how success will be estimated.
- For the metric, consider the higher-level goal, problem type, and label structure.
- For the protocol, begin with data availability.
What do you think happens?
A project has little data and requires highly accurate model evaluation. Which protocol is the source's recommended choice?
Reveal answer
Answer: Iterated K-fold validation
Iterated K-fold validation repeats K-fold cross-validation multiple times and is recommended when highly accurate evaluation is needed with little data.
Key Takeaways
- Define observable success before choosing or training a model.
- Choose a metric that aligns with the higher-level goal, the problem type, and the label structure.
- Keep the metric separate from the evaluation protocol: the metric defines performance, while the protocol describes how performance is assessed.
- Use hold-out validation when data is plentiful, K-fold cross-validation when samples are limited, and iterated K-fold validation when little data must support highly accurate evaluation.
- A meaningful metric and a suitable evaluation protocol are both necessary for a well-supported progress estimate.
Key Takeaways
- A machine learning project needs an observable definition of success before progress can be measured.
- Metric choice depends on the higher-level goal, problem type, and label structure.
- An evaluation protocol is the procedure used to assess the chosen metric.
- Hold-out validation suits plentiful data, K-fold cross-validation suits limited data, and iterated K-fold validation suits highly accurate evaluation with little data.