Concepts / Machine learning model evaluation

Machine learning model evaluation

Model evaluation starts with a deliberate division of data into training, validation, and test sets.

  • Programming

Evaluation Starts with Data Roles

Model evaluation begins before you calculate a score. First, decide how the available data will be divided and what role each part will play. The central division is into training, validation, and test sets. Training data supports model training, validation data supports the evaluation process used while working with the model, and test data provides the final evaluation role. Keeping these roles distinct makes the evaluation plan deliberate rather than accidental.

dividedividedivideAvailable dataTraining setmodel trainingValidation setevaluation processTest setfinal evaluation
Which data records are used for training, model selection, and final unbiased evaluation, and how are these roles kept separate?

A Worked Data Division

Assigning roles before scoring

Imagine a dataset that must be organized for a model evaluation process. How can its records be assigned distinct roles?

Start with the available data: Treat the complete collection of available records as the source from which the evaluation roles will be created.

Assign training data: Place one portion in the training role so it is available for model training.

Assign validation data: Place another portion in the validation role for the evaluation process.

Reserve test data: Keep a separate portion in the test role for the final evaluation.

The important result is not a particular numeric split. It is a deliberate separation of roles before calculating a score.

The Single Hold-Out Split

Simple hold-out validation is the basic splitting recipe. Instead of creating several fold-based arrangements, it makes one division: one portion is used for training and one portion is held out for validation. This can be enough in some situations. Its defining feature is that the training and validation arrangement is a single split rather than a collection of fold-based arrangements.

single splitsingle splitAvailable dataTraining portionHeld-out portionvalidation
How does a single split reserve one portion of the dataset for validation while the rest is used for training?

A team might use simple hold-out validation when a single training portion and a single held-out validation portion fit the situation. The method is recognized by its one training-validation split.

K-Fold Rotation

K-fold validation changes the arrangement from one split to multiple folds. The dataset is organized into several folds, and the method forms multiple training and validation arrangements from those folds. In each arrangement, one fold serves as the validation portion while the remaining folds serve as the training portion. The validation role therefore rotates across the folds.

next arrangementnext arrangementFold AvalidationFold BvalidationFold CvalidationFold BtrainingFold AtrainingFold CtrainingFold Ctraining
How does each fold take a turn serving as validation data while the remaining folds are used for training?

Reading a three-fold arrangement

Suppose a dataset has three folds named A, B, and C. What does the fold rotation represent?

First arrangement: Fold A is the validation portion, while folds B and C form the training portion.

Second arrangement: Fold B becomes the validation portion, while folds A and C form the training portion.

Third arrangement: Fold C becomes the validation portion, while folds A and B form the training portion.

Each fold takes a turn in the validation role, while the other folds take the training role.

Shuffling and Repeating Folds

Iterated K-fold validation with shuffling extends K-fold validation in two ways. It adds shuffling to the dataset arrangement, and it repeats the K-fold approach. Repetition with different fold assignments creates additional fold-based training and validation arrangements. The defining distinction is therefore not merely that folds are used, but that shuffling and repetition are added to the K-fold process.

organize into foldsshuffle and repeatorganize into folds againDatasetinitial arrangementK-fold arrangementfold assignment 1K-fold arrangementfold assignment 2Shuffled datasetnew arrangement
What changes when the dataset is shuffled and K-fold validation is repeated with different fold assignments?
RecipeData arrangementDefining featureWhen the source suggests considering it
Simple hold-out validationOne training portion and one held-out validation portionA single splitWhen a single split is enough
K-fold validationSeveral fold-based training and validation arrangementsValidation rotates across foldsWhen a single split is not the preferred arrangement
Iterated K-fold validation with shufflingRepeated fold-based arrangements after shufflingShuffling and repetition extend K-fold validationWhen little data is available and a more advanced splitting method is useful

Choosing with Limited Data

The source material gives a practical starting point: a single split may be enough in some situations, but more advanced splitting methods can be useful when little data is available. Use the method names to describe the arrangement you need. Choose simple hold-out validation when one training-validation split is appropriate. Choose K-fold validation when you want several fold-based training and validation arrangements. Choose iterated K-fold validation with shuffling when you want K-fold validation extended by shuffling and repetition.

may be enoughcan be usefulcan be usefulSimple hold-outsingle splitLimited dataconsider advanced methodsK-foldmultiple arrangementsIterated K-foldshuffle and repeat
How do simple hold-out, K-fold, and iterated K-fold validation differ in their data arrangements and their use when data is limited?

Common Recognition Errors

  • Treating model evaluation as a scoring step only.

    The source defines the division of available data as the starting point of evaluation.

    Fix: Describe the data roles and splitting recipe before discussing the score.

  • Calling every split K-fold validation.

    Simple hold-out validation is the basic single-split recipe, while K-fold validation uses multiple folds to form several arrangements.

    Fix: Use simple hold-out for one split and K-fold for multiple fold-based arrangements.

  • Forgetting what iterated K-fold adds.

    Iterated K-fold validation with shuffling specifically adds shuffling and repetition to K-fold validation.

    Fix: Mention both additions when identifying the iterated method.

  • Assuming limited data automatically dictates one named recipe.

    The source says advanced methods can be useful when little data is available, while also noting that a single split may be enough in some situations.

    Fix: Consider whether a single split is enough, then distinguish K-fold from iterated K-fold according to the arrangement required.

Check Your Choice

MEDIUM

A dataset is limited. You first consider one training portion and one held-out validation portion. Then you consider several fold-based arrangements. Finally, you consider shuffling the dataset and repeating the fold-based process. Name the three evaluation recipes in the order described and state the defining change from one recipe to the next.

Hints
  • The first recipe is the basic splitting recipe.
  • The second recipe uses multiple folds.
  • The third recipe adds shuffling and repetition to the second recipe.

What do you think happens?

Which description identifies iterated K-fold validation with shuffling?

  • One training portion and one held-out validation portion
  • Several fold-based training and validation arrangements
  • Shuffling the dataset and repeating the K-fold approach
Reveal answer

Answer: Shuffling the dataset and repeating the K-fold approach

The source describes iterated K-fold validation with shuffling as K-fold validation extended by shuffling and repetition.

Key Takeaways

  1. Model evaluation begins by deliberately dividing available data into training, validation, and test roles.
  2. Simple hold-out validation uses one training-validation split.
  3. K-fold validation uses multiple folds to form several training and validation arrangements, with the validation role rotating across folds.
  4. Iterated K-fold validation with shuffling adds shuffling and repetition to the K-fold approach.
  5. When data is limited, a single split may be enough in some situations, while more advanced splitting methods can also be useful.

Key Takeaways

  • Evaluation starts with a deliberate data division, not with score calculation.
  • Training, validation, and test sets have separate roles in the evaluation process.
  • Simple hold-out uses one split; K-fold uses several fold-based arrangements.
  • Iterated K-fold validation with shuffling adds shuffling and repetition to K-fold validation.
  • For limited data, consider whether a single split is enough or whether an advanced fold-based method is useful.