Loss Functions and Optimization
A machine learning project needs an observable definition of success before progress can be measured.
Start with a Target
A machine learning project cannot improve toward an undefined target. Before choosing a model or deciding how to train it, define what success means and how success will be observed. This observable definition becomes the measure of success used to evaluate model performance.
Do not choose a metric only because it is convenient to calculate. The metric should connect directly to the higher-level goal, such as the success of a business.
Metric Choice and Loss Choice
The measure of success is the metric used to evaluate model performance. Metric choice depends on the problem type. The source distinguishes balanced-classification, class-imbalanced, ranking, and multilabel problems as different problem categories, so the appropriate success measure must be considered in relation to the category rather than selected automatically.
A Generated Decision Example
Choosing what to measure first
A team is building a model for a machine learning problem. It can calculate a convenient performance number, but it has not yet connected that number to the project's higher-level goal.
Identify the goal: The team first states what success means for the larger project rather than starting with a convenient calculation.
Make success observable: The team defines a measure of success that can be used to evaluate model performance.
Classify the problem: The team considers whether the problem is balanced-classification, class-imbalanced, ranking, or multilabel, because metric choice depends on problem type.
Connect the metric to training: The selected measure guides the choice of loss function and the way progress is assessed.
The team has an observable target aligned with the higher-level goal instead of an unexplained convenient number.
This example illustrates the central dependency: optimization decisions are meaningful only after the project has defined what it is trying to improve. A carefully executed evaluation of the wrong metric is still progress toward the wrong target.
What an Evaluation Protocol Does
A metric produces the notion of performance that matters. An evaluation protocol establishes the procedure used to assess that performance. The protocol determines how available data is organized for validation and how the resulting performance estimate is obtained.
Three Ways to Use the Data
| Protocol | How it uses the data | When the source recommends it |
|---|---|---|
| Hold-out validation | Sets aside part of the available data for validation | When plenty of data is available |
| K-fold cross-validation | Divides the data into folds; each fold can serve as the evaluation portion while the other folds provide the remaining data in that cycle | When there are too few samples for hold-out validation to be reliable |
| Iterated K-fold validation | Repeats K-fold cross-validation multiple times | When highly accurate evaluation is needed with little data |
Hold-out validation is the simple choice when plenty of data is available. Setting aside a validation portion is less likely to make the remaining training data too small for the task. When data is limited, K-fold cross-validation uses multiple evaluation cycles so that each fold can serve as the evaluation portion while the other folds provide the remaining data for that cycle.
Iterated K-fold validation changes K-fold cross-validation through repetition. Instead of carrying out K-fold cross-validation once, the procedure is carried out multiple times. This makes it the source's choice for highly accurate model evaluation when little data is available.
Trace the Protocol Choice
What do you think happens?
A project has plenty of data. Which protocol is the source's straightforward choice?
Reveal answer
Answer: Hold-out validation
The source describes hold-out validation as the simple choice when plenty of data is available because reserving a portion for validation is less likely to make the remaining training data too small.
What do you think happens?
A project has little data and requires highly accurate evaluation. Which protocol best matches the source's recommendation?
Reveal answer
Answer: Iterated K-fold validation
The source recommends iterated K-fold validation when highly accurate model evaluation is needed with little data. Its defining feature is that K-fold cross-validation is repeated multiple times.
Selecting a validation protocol
Compare three generated project situations: one with plentiful data, one with too few samples for a reliable hold-out estimate, and one with little data where highly accurate evaluation is especially important.
Plentiful data: Choose hold-out validation and set aside part of the available data for validation.
Too few samples for hold-out: Choose K-fold cross-validation so that each fold can serve as the evaluation portion across the procedure.
Little data plus high accuracy requirement: Choose iterated K-fold validation, which repeats the K-fold procedure multiple times.
Data availability is the main deciding factor, with repeated K-fold evaluation reserved for highly accurate evaluation when little data is available.
Mistakes in Evaluation Design
Choosing a metric because it is convenient to calculate
The metric may measure performance carefully while failing to represent the success that matters for the larger project.
Fix:
Define the higher-level goal first, then choose an observable measure of success that connects directly to it.Treating the metric and the evaluation protocol as the same decision
A strong metric without a suitable protocol leaves the progress estimate poorly supported.
Fix:
Choose the metric for what performance means and the protocol for how that performance will be assessed.Using hold-out validation automatically when data is limited
The source recommends K-fold cross-validation when there are too few samples for hold-out validation to be reliable.
Fix:
Use K-fold cross-validation for limited samples, or iterated K-fold validation when little data must support highly accurate evaluation.Confusing K-fold validation with iterated K-fold validation
Iteration means repeating K-fold cross-validation multiple times.
Fix:
Reserve the term iterated K-fold validation for the repeated procedure.
Practice the Selection
A machine learning project has a clearly stated higher-level goal, but the team has not yet defined how success will be observed. The dataset is small, and the team wants highly accurate evaluation. State the next two decisions the team should make: first, what kind of target must be defined; second, which evaluation protocol best fits the data situation.
Hints
- The target should be an observable measure of success aligned with the higher-level goal.
- Little data combined with a need for highly accurate evaluation points to the repeated form of K-fold validation.
The answer is: define a measure of success that directly connects to the higher-level goal, then use iterated K-fold validation because the project has little data and requires highly accurate evaluation.
Key Takeaways
- A machine learning project needs an observable definition of success before progress can be measured.
- The metric should align with the higher-level goal, and metric choice depends on the problem type.
- A metric defines what performance means, while an evaluation protocol defines how that performance is assessed.
- Use hold-out validation when data is plentiful, K-fold cross-validation when samples are too limited for reliable hold-out validation, and iterated K-fold validation when little data must support highly accurate evaluation.
- Iterated K-fold validation is K-fold cross-validation repeated multiple times.