Concepts / Model evaluation on unseen data

Model evaluation on unseen data

Optimization and generalization measure different kinds of model performance.

  • Programming

Why Familiar Success Can Mislead

A model can appear to improve while it is learning less useful behavior. Its performance on the data used for training may continue to get better, yet its performance on data it has not seen may stop improving or begin to decline. This is the central difficulty behind overfitting: success on familiar examples is not the same as reliable performance on new examples.

Optimization measures how well the model is improving on the training data. Generalization measures how reliably that learned behavior works on data the model has not seen.

improves onreveals performance onTraining dataFamiliar examplesOptimizationPerformance improvesUnseen dataNew examplesGeneralizationMay stall or decline
How can training performance improve while performance on unseen data gets worse?

Three Stages of Training

Training performance and validation performance often move together at first. Reducing the model's loss on training examples is commonly accompanied by lower loss on unseen examples. This early agreement is useful, but it does not continue indefinitely.

training continuestraining continuesUnderfittingBoth can improveStrong validationBest observed pointOverfittingTraining improves;validation worsens
What changes in training and validation performance as a model moves through training?

Following a topic-classification model

A model is trained to classify topics. What do three observations during training tell us?

Beginning: The model has not captured all the relevant patterns in the training data. It is underfit, so useful progress is still available.

Middle: Training performance improves and validation performance also improves. The model is learning behavior that is useful beyond the familiar training examples.

Later: Training performance continues to improve, but validation metrics stop improving and then worsen. Optimization is continuing, while generalization is no longer getting better.

The model has moved from underfitting toward overfitting. The strongest validation performance occurred before the later training progress on the training set.

Underfitting does not simply mean that a model performs poorly for no reason. It means the model has not yet represented all the relevant patterns in its training data, so further useful progress is still available.

Reading Validation Performance

Validation performance helps reveal when continued training has stopped improving generalization. The training result tells you how the model is doing on examples it has used. The validation result gives a separate view of how dependable the learned behavior is on examples it has not seen.

assigned for learningheld out for evaluationtrainsevaluatesproducesAvailable examplesTraining dataUsed during trainingModelLearns from training dataValidation dataKept unseen during trainingValidationperformanceEvidence aboutgeneralization
How does data move from training to evaluation, and why must evaluation data remain unseen during training?

The key comparison is not whether training performance is improving in isolation. Ask whether validation performance is improving as well. If both improve, the model may still be learning useful patterns. If training performance improves while validation performance stalls or declines, the model is becoming increasingly tailored to its training set.

Useful Patterns and Memorized Details

A useful pattern captures information that helps the model perform on new examples. A training-specific pattern belongs to the particular training data and does not provide reliable guidance outside it. Overfitting begins when the model starts learning details that may be misleading or irrelevant beyond the training set.

transfers tofails to guideRelevant patternUseful on new examplesNew examplesReliable guidanceTraining-specificdetailMisleading outside trainingdataUnreliable resultDoes not transfer reliably
Which patterns learned from training data remain useful on new examples, and which are merely specific to the training set?

In movie-review prediction, topic classification, or house-price regression, a model may learn patterns that are genuinely useful across examples. It may also learn details specific to the training data that do not help with new reviews, new documents, or new houses. The task changes, but the basic overfitting issue remains the same.

The distinction cannot be decided from training performance alone. A training-specific pattern can make the model look more successful on familiar examples even while making its behavior less dependable on unseen data. Validation performance is therefore evidence about whether learned behavior transfers.

Reducing Overfitting

The strongest general remedy for overfitting is to obtain more training data. More examples give the model more information from which to learn and make it more likely to develop patterns that generalize.

When obtaining more data is not possible, the next approaches change how much information the model can store or place constraints on the information it is allowed to store. Limiting what the model can retain makes unrestricted memorization less available. The optimization process is then pushed toward the most prominent patterns, which have a better chance of generalizing.

supports learningpushes optimization towardhelps identifyMore training dataMore examples to learn fromMore reliablepatternsBetter chance of transferLimit retainedinformationLess unrestrictedmemorizationMonitor validationNotice when generalizationstalls
How do more data, limits on retained information, and validation-guided training change the tendency to overfit?
  • Compare training performance with validation performance rather than relying on training performance alone.
  • Treat continued training as useful only while validation performance is also improving.
  • Prefer more training data when it can be obtained.
  • When more data is unavailable, limit how much information the model can retain so unrestricted memorization is less available.

Mistakes to Avoid

  • Assuming that better training performance always means a better model

    Optimization is continuing, but generalization is declining.

    Fix: Check whether validation performance improves alongside training performance.

  • Calling every poorly performing model overfit

    This describes underfitting, where useful progress is still available.

    Fix: Distinguish a model that has learned too little from one that has learned training-specific details.

  • Treating familiar examples as evidence of reliable performance on new data

    Success on familiar examples does not establish generalization.

    Fix: Use validation performance to look for evidence that learned behavior transfers.

  • Continuing training without watching for a validation decline

    The model is moving into overfitting.

    Fix: Treat the validation trend as a signal that further training may no longer improve generalization.

Check Your Reasoning

MEDIUM

A model's performance on training data improves throughout training. Its validation performance improves at first, then stops improving, and finally worsens. Identify the training stage at the beginning and the stage at the end. Then explain what happened to optimization and generalization.

Hints
  • At the beginning, ask whether the model has captured all relevant patterns.
  • At the end, compare the direction of training performance with the direction of validation performance.
  • Use optimization for progress on training data and generalization for reliability on unseen data.

Reasoning through the trend

Interpret a model whose training performance improves continuously, while validation performance first improves and later declines.

Initial stage: The model is underfit because it has not yet captured all the relevant patterns in the training data.

Improving stage: Training and validation performance improve together, so optimization is accompanied by better generalization.

Declining validation stage: Training performance continues to improve, but validation performance worsens. The model is learning details specific to the training data.

The model has changed from underfitting to overfitting. Optimization continues, but generalization has stopped improving.

Key Takeaways

  1. Optimization and generalization measure different kinds of model performance.
  2. Underfitting means that relevant patterns in the training data have not yet been captured.
  3. Overfitting begins when training-specific details improve familiar-data performance without transferring reliably to new data.
  4. Validation performance helps identify when continued training has stopped improving generalization.
  5. More training data is the strongest general remedy; when that is unavailable, limiting retained information can reduce unrestricted memorization.

Key Takeaways

  • Optimization concerns performance on training data, while generalization concerns performance on unseen data.
  • Training and validation performance often improve together at first, but validation performance can later stall or worsen.
  • That divergence marks the movement from useful learning toward overfitting.
  • More data and limits on retained information can reduce the tendency to memorize training-specific details.