Concepts / Overfitting

Overfitting

Data augmentation applies transformations to existing data.

  • Programming

The Familiar-Data Trap

A model can appear to improve while becoming less dependable. Its performance on the examples used for training may continue to get better, while its performance on data it has not seen stops improving or begins to decline. This is the central difficulty of overfitting: success on familiar examples is not the same as reliable performance on new examples.

Overfitting occurs when a model learns patterns that are specific to the training data and do not transfer reliably to new data.

The important question is not only whether the model performs well on the data it has already seen. The question is whether the learned behavior remains useful on data it has not seen.

Training Through Three Phases

Training often begins with underfitting. At this point, the model has not captured all the relevant patterns in the training data, so useful progress is still available. As training continues, performance on both the training examples and unseen examples can improve. After a certain point, however, validation performance may reach its strongest level. Continued training can still improve the training result while validation performance stalls and then worsens. That is the transition into overfitting.

training continuescontinued trainingUnderfittingRelevant patterns remainunlearnedGeneralizationTraining and validationimproveOverfittingTraining improves;validation worsens
What changes in training and validation performance as a model moves from underfitting to useful generalization and then to overfitting?

Reading the Training Story

A model is observed at three points during training. Early on, it performs poorly on both training and validation data. Later, both results improve. After still more training, the training result improves again, but the validation result declines. Identify the phase at each point.

First point: The model has not captured all relevant patterns in its training data. This is underfitting.

Second point: Training and validation performance improve together. The model is learning behavior that is useful beyond the familiar training examples.

Third point: Training performance continues to improve, but validation performance declines. The model is becoming increasingly tailored to the training set, which indicates overfitting.

The sequence is underfitting, useful generalization, and then overfitting.

Optimization and Generalization

Optimization and generalization describe different kinds of model performance. Optimization concerns improvement on the training examples. Generalization concerns how reliably the model performs on data it has not seen. These measures often move together at the beginning of training, but they do not have to continue moving together indefinitely.

may diverge fromOptimizationTraining performanceimprovesGeneralizationValidation performancestalls or worsens
How can a model's training performance improve while its performance on unseen data gets worse?

Validation performance helps reveal when continued training has stopped improving generalization. If training performance keeps improving but validation metrics stop improving, the model is no longer gaining dependable performance on new data. If validation metrics then worsen, the model is moving further into overfitting.

What do you think happens?

A model's training performance improves for several more training steps, but its validation performance has already begun to decline. What is the best interpretation?

  • The model is necessarily generalizing better
  • Optimization is continuing, but generalization is worsening
  • The model has returned to underfitting
  • Training and validation now measure exactly the same thing
Reveal answer

Answer: Optimization is continuing, but generalization is worsening

Improvement on training examples does not guarantee improvement on unseen examples. Continued training can make the model more tailored to the training set while validation performance declines.

Patterns That Transfer

A useful pattern captures regularities that help the model perform on new examples. A training-specific pattern is tied to details of the training data and does not provide reliable guidance outside that data. Overfitting begins when the model starts learning these training-specific details instead of concentrating on patterns that are more dependable for new data.

helps predictdoes not reliably helpUseful patternsRemain useful on new dataTraining-specificpatternsDo not transfer reliablyNew examples
Which patterns learned from training data remain useful on new examples, and which patterns only help memorize the training set?

In topic classification, early improvements on training examples are encouraging only when validation performance improves as well. If validation metrics later stop improving while training performance continues to improve, the two measurements indicate that optimization is continuing but generalization is no longer improving.

Augmenting the Training Collection

Data augmentation is a technique that applies transformations to existing data to artificially increase the size of a dataset.

The original data is the starting point for the augmented collection. Transformations are applied to those existing examples, producing additional training examples. The result is an artificially larger dataset. Data augmentation is used to mitigate overfitting because a model can otherwise be trained on a dataset that does not contain enough variation for the learning task.

applyapplyproducesproducesOriginal exampleStarting dataTransformation AApplied to existing dataAdditional example ATransformed dataTransformation BApplied to existing dataAdditional example BTransformed data
How does one original example become multiple transformed examples in an augmented dataset?
transformation producesOriginal exampleStarting pointOriginal exampleRetained as the startingdataAdditional examplesCreated throughtransformations
What is the relationship between an original example and the augmented examples created from it?

Separating the Roles of the Data

A dataset contains original examples. Transformations are applied to those examples, producing additional examples for training. Which data is original, and which data is transformed?

Identify the starting point: The examples that existed before the transformations are the original data.

Identify the new training examples: The examples produced by applying transformations are the additional transformed data.

Identify the purpose: Together, the original and additional examples form an artificially larger training collection intended to provide more variation and help mitigate overfitting.

Original data starts the process; transformed data expands the training collection.

Reducing Overfitting

The strongest general remedy for overfitting is to obtain more training data. More data gives the model more examples from which to learn and makes it more likely that the model develops patterns that generalize. Data augmentation is one way to artificially increase the number of training examples when the existing collection does not contain enough variation.

When obtaining more data is not possible, other approaches change how much information the model can store or place constraints on the information it is allowed to store. Limiting what the model can retain makes unrestricted memorization less available. The optimization process is then pushed toward the most prominent patterns, which have a better chance of generalizing.

  • Obtain more training data when possible.
  • Use data augmentation to apply transformations to existing data and create additional training examples.
  • Limit how much information the model can retain when more data is unavailable.
  • Place constraints on the information the model is allowed to store so unrestricted memorization is less available.
  • Monitor validation performance to identify when continued training stops improving generalization.

Common Misreadings

  • Treating improved training performance as proof of improved generalization.

    Optimization and generalization measure different kinds of performance.

    Fix: Compare training behavior with validation behavior. Continued training improvement can coexist with worsening performance on unseen data.

  • Calling every poorly performing model overfit.

    This describes underfitting, not overfitting.

    Fix: Use underfitting when useful patterns remain unlearned. Use overfitting when the model learns training-specific patterns that do not transfer reliably.

  • Treating transformed examples as unrelated to the original data.

    Data augmentation applies transformations to existing data, and the original data is the starting point for the augmented collection.

    Fix: Recognize transformed examples as additional training examples produced from existing data.

  • Assuming that more training always improves the model's usefulness on new data.

    Further training can make the model increasingly tailored to the training set.

    Fix: Watch validation performance for signs that generalization has stalled or worsened.

Check Your Understanding

MEDIUM

A model is trained for a longer period. Its training performance improves throughout the process. Its validation performance first improves, then stops improving, and finally becomes worse. Explain what happened, identify the phase in which overfitting began, and name two strategies that could help reduce the problem.

Hints
  • Compare optimization with generalization.
  • The turning point is where validation performance stops improving even though training performance continues to improve.
  • Possible strategies include obtaining more data, using data augmentation, limiting retained information, or placing constraints on what the model can store.
EASY

An existing dataset does not contain enough variation for a learning task. Explain how data augmentation changes the training collection and why this can help with overfitting.

Hints
  • Start with the role of the original data.
  • Explain what transformations produce.
  • Connect the artificially larger collection to the model's tendency to overfit.

Key Takeaways

  1. Overfitting happens when a model learns training-specific patterns that do not transfer reliably to new data.
  2. Optimization can continue improving training performance after generalization has stopped improving.
  3. Underfitting can progress into useful generalization and then into overfitting as training continues.
  4. Data augmentation applies transformations to existing data to create an artificially larger collection of training examples.
  5. More data, data augmentation, limits on retained information, constraints on stored information, and validation monitoring can help reduce overfitting.

Key Takeaways

  • Overfitting is the learning of training-specific patterns that do not reliably help on new data.
  • Training performance and validation performance may improve together at first, but they can diverge as training continues.
  • Data augmentation transforms existing examples to create additional training data and help mitigate overfitting.
  • Validation performance is an important signal for identifying when a model has stopped generalizing better.
  • Obtaining more data and limiting unrestricted memorization are practical ways to reduce overfitting.