Machine learning model evaluation
Model evaluation starts with a deliberate division of data into training, validation, and test sets.
Evaluation Starts with Data Roles
Model evaluation begins before you calculate a score. First, decide how the available data will be divided and what role each part will play. The central division is into training, validation, and test sets. Training data supports model training, validation data supports the evaluation process used while working with the model, and test data provides the final evaluation role. Keeping these roles distinct makes the evaluation plan deliberate rather than accidental.
A Worked Data Division
Assigning roles before scoring
Imagine a dataset that must be organized for a model evaluation process. How can its records be assigned distinct roles?
Start with the available data: Treat the complete collection of available records as the source from which the evaluation roles will be created.
Assign training data: Place one portion in the training role so it is available for model training.
Assign validation data: Place another portion in the validation role for the evaluation process.
Reserve test data: Keep a separate portion in the test role for the final evaluation.
The important result is not a particular numeric split. It is a deliberate separation of roles before calculating a score.
The Single Hold-Out Split
Simple hold-out validation is the basic splitting recipe. Instead of creating several fold-based arrangements, it makes one division: one portion is used for training and one portion is held out for validation. This can be enough in some situations. Its defining feature is that the training and validation arrangement is a single split rather than a collection of fold-based arrangements.
A team might use simple hold-out validation when a single training portion and a single held-out validation portion fit the situation. The method is recognized by its one training-validation split.
K-Fold Rotation
K-fold validation changes the arrangement from one split to multiple folds. The dataset is organized into several folds, and the method forms multiple training and validation arrangements from those folds. In each arrangement, one fold serves as the validation portion while the remaining folds serve as the training portion. The validation role therefore rotates across the folds.
Reading a three-fold arrangement
Suppose a dataset has three folds named A, B, and C. What does the fold rotation represent?
First arrangement: Fold A is the validation portion, while folds B and C form the training portion.
Second arrangement: Fold B becomes the validation portion, while folds A and C form the training portion.
Third arrangement: Fold C becomes the validation portion, while folds A and B form the training portion.
Each fold takes a turn in the validation role, while the other folds take the training role.
Shuffling and Repeating Folds
Iterated K-fold validation with shuffling extends K-fold validation in two ways. It adds shuffling to the dataset arrangement, and it repeats the K-fold approach. Repetition with different fold assignments creates additional fold-based training and validation arrangements. The defining distinction is therefore not merely that folds are used, but that shuffling and repetition are added to the K-fold process.
| Recipe | Data arrangement | Defining feature | When the source suggests considering it |
|---|---|---|---|
| Simple hold-out validation | One training portion and one held-out validation portion | A single split | When a single split is enough |
| K-fold validation | Several fold-based training and validation arrangements | Validation rotates across folds | When a single split is not the preferred arrangement |
| Iterated K-fold validation with shuffling | Repeated fold-based arrangements after shuffling | Shuffling and repetition extend K-fold validation | When little data is available and a more advanced splitting method is useful |
Choosing with Limited Data
The source material gives a practical starting point: a single split may be enough in some situations, but more advanced splitting methods can be useful when little data is available. Use the method names to describe the arrangement you need. Choose simple hold-out validation when one training-validation split is appropriate. Choose K-fold validation when you want several fold-based training and validation arrangements. Choose iterated K-fold validation with shuffling when you want K-fold validation extended by shuffling and repetition.
Common Recognition Errors
Treating model evaluation as a scoring step only.
The source defines the division of available data as the starting point of evaluation.
Fix:
Describe the data roles and splitting recipe before discussing the score.Calling every split K-fold validation.
Simple hold-out validation is the basic single-split recipe, while K-fold validation uses multiple folds to form several arrangements.
Fix:
Use simple hold-out for one split and K-fold for multiple fold-based arrangements.Forgetting what iterated K-fold adds.
Iterated K-fold validation with shuffling specifically adds shuffling and repetition to K-fold validation.
Fix:
Mention both additions when identifying the iterated method.Assuming limited data automatically dictates one named recipe.
The source says advanced methods can be useful when little data is available, while also noting that a single split may be enough in some situations.
Fix:
Consider whether a single split is enough, then distinguish K-fold from iterated K-fold according to the arrangement required.
Check Your Choice
A dataset is limited. You first consider one training portion and one held-out validation portion. Then you consider several fold-based arrangements. Finally, you consider shuffling the dataset and repeating the fold-based process. Name the three evaluation recipes in the order described and state the defining change from one recipe to the next.
Hints
- The first recipe is the basic splitting recipe.
- The second recipe uses multiple folds.
- The third recipe adds shuffling and repetition to the second recipe.
What do you think happens?
Which description identifies iterated K-fold validation with shuffling?
Reveal answer
Answer: Shuffling the dataset and repeating the K-fold approach
The source describes iterated K-fold validation with shuffling as K-fold validation extended by shuffling and repetition.
Key Takeaways
- Model evaluation begins by deliberately dividing available data into training, validation, and test roles.
- Simple hold-out validation uses one training-validation split.
- K-fold validation uses multiple folds to form several training and validation arrangements, with the validation role rotating across folds.
- Iterated K-fold validation with shuffling adds shuffling and repetition to the K-fold approach.
- When data is limited, a single split may be enough in some situations, while more advanced splitting methods can also be useful.
Key Takeaways
- Evaluation starts with a deliberate data division, not with score calculation.
- Training, validation, and test sets have separate roles in the evaluation process.
- Simple hold-out uses one split; K-fold uses several fold-based arrangements.
- Iterated K-fold validation with shuffling adds shuffling and repetition to K-fold validation.
- For limited data, consider whether a single split is enough or whether an advanced fold-based method is useful.