Concepts / Training and validation performance

Training and validation performance

Optimization and generalization measure different kinds of model performance.

  • Programming

Why Two Performances Matter

A model can appear to improve while becoming less useful. Its performance on the examples used during training may continue to get better, while its performance on data it has not seen stops improving or begins to decline. To understand whether training is producing reliable behavior, compare training performance with validation performance.

Optimization is the process of improving a model's performance on its training examples. Generalization is the model's ability to perform reliably on new data. These measure different kinds of performance.

Tracing the Training Stages

training continuesvalidation improvestraining continuesgeneralization declinesBeginningtraining: limited;validation: limitedImprovingtraining: better;validation: betterStrongest validationvalidation reaches its bestpointContinued trainingtraining: better;validation: stallsOverfittingtraining: better;validation: worse
What changes in training and validation performance as the model moves from underfitting to good fit and then to overfitting?

At the beginning, the model has not captured all the relevant patterns in the training data. This is underfitting. As training proceeds, training performance and validation performance often improve together because the model is learning useful patterns. Later, validation performance can reach its strongest point even though training performance still has room to improve.

Optimization Versus Generalization

measures performance onreveals performance onOptimizationtraining performanceimprovesTraining dataexamples used duringtrainingGeneralizationnew-data performanceValidation datadata the model has not seen
How can a model continue improving on the training data while becoming worse on unseen validation data?

Optimization and generalization can agree for a while, but they are not the same goal. Optimization focuses on the data used to train the model. Generalization focuses on whether the learned behavior transfers to data the model has not seen. Early in training, reducing the model's loss on training examples is commonly accompanied by better performance on unseen examples. That agreement may later end.

A topic-classification model

A model is trained to classify the topics of text examples. Its training performance keeps improving after validation performance has reached its strongest point.

Early training: The model has not yet represented all relevant patterns. Improvements on the training examples are useful when validation performance improves as well.

Best validation point: Validation performance reaches its strongest point. The model is showing its most dependable performance on data it has not seen.

Additional training: Training performance continues to improve, but validation performance stalls and then worsens. Optimization is continuing, while generalization is no longer improving.

Interpretation: The model is moving into overfitting because it is learning details specific to the training data rather than behavior that transfers reliably.

Training performance alone does not show whether the model is becoming more dependable. Validation performance reveals when additional training has stopped helping generalization.

Useful and Training-Specific Patterns

transfers reliablyimproves performance onUseful patternhelps training and new dataNew datareliable performanceTraining-specificpatternhelps training onlyTraining dataimproved familiar examples
Which patterns learned from the training set also help on new data, and which patterns only improve training performance?

Useful learning captures patterns that improve performance beyond the familiar training examples. Overfitting begins when the model starts learning details that belong specifically to the training data but do not provide reliable guidance for new data. Those details may be misleading or irrelevant outside the training set.

The same issue can appear in different tasks, including movie-review prediction, topic classification, and house-price regression. In each case, the central question is whether improved performance on training examples is accompanied by improved performance on unseen examples.

The Underfitting-to-Overfitting Transition

useful patterns are learnedtraining continuesvalidation declinesUnderfittingrelevant patterns remain tolearnStrong validationgeneralization is strongestValidation stallsadditional training becomesquestionableOverfittingtraining improves;validation worsens
How does the model's behavior change across training stages, and where does additional training start harming generalization?

Underfitting is not simply poor performance with no explanation. It means that the model has not yet represented all the relevant patterns in the training data, so useful progress is still available. Overfitting is a later problem: the model has begun learning training-specific details that do not transfer reliably.

What do you think happens?

Validation performance has stopped improving, but training performance is still improving. What is the most likely interpretation?

  • The model is still improving its generalization
  • The model is moving into overfitting
  • The model has returned to underfitting
Reveal answer

Answer: The model is moving into overfitting

The training results show continued optimization, but the validation results show that generalization has stalled or begun to decline.

Reducing Overfitting

can contribute tosupports learning frompushes optimization towardTraining-specificdetailsmore available to retainTraining-validationgapcan grow as overfittingdevelopsMore training datamore examples availableInformationconstraintsless unrestrictedmemorizationProminent patternsbetter chance ofgeneralizing
How do practical strategies change the gap between training and validation performance?

The strongest general remedy for overfitting is to obtain more training data. More examples give the model more information from which to learn and make it more likely to develop patterns that generalize.

When obtaining more data is not possible, use approaches that change how much information the model can store or place constraints on the information it is allowed to store. Limiting what the model can retain makes unrestricted memorization less available. The optimization process is then pushed toward prominent patterns that have a better chance of generalizing.

Compare training and validation performance throughout training. Treat continued training as questionable when training performance keeps improving but validation performance stalls or worsens. Prefer more training data when it is available; otherwise, consider approaches that limit what the model can retain.

Mistakes in Reading Performance

  • Treating better training performance as proof that the model is becoming more useful.

    Optimization on familiar examples can continue while generalization to unseen data gets worse.

    Fix: Check whether validation performance is improving as well.

  • Calling every poorly performing model overfit.

    This describes underfitting: useful progress is still available.

    Fix: Use underfitting when the model has not represented all relevant training patterns, and overfitting when it learns training-specific patterns that fail to transfer.

  • Assuming training and validation performance must move together forever.

    The early agreement between the measurements does not continue indefinitely.

    Fix: Continue comparing the measurements after the early stage because validation performance can stall and then worsen.

  • Focusing only on memorizing more training information as a solution.

    Unrestricted memorization can make training performance look better without improving performance on new data.

    Fix: Obtain more training data when possible, or limit the information the model can retain.

Check Your Understanding

MEDIUM

A model's training performance improves during three successive stages. Validation performance improves during the first two stages, reaches its strongest point at the second stage, and worsens during the third stage. Identify the likely training state at each stage and explain why the third stage should not be judged successful from training performance alone.

Hints
  • At the beginning, ask whether the model has captured all relevant patterns.
  • At the strongest validation point, ask whether generalization is at its best.
  • At the final stage, compare the direction of training performance with the direction of validation performance.

Interpreting the three stages

Use the described performance changes to identify the model's behavior.

Stage one: The model is likely underfitting because useful patterns remain to be learned.

Stage two: The model has reached its strongest validation performance, so generalization is currently strongest.

Stage three: The model is moving into overfitting because training performance improves while validation performance worsens.

The model's training history moves from underfitting through its strongest validation point and then toward overfitting.

Key Takeaways

  1. Optimization measures improvement on training examples, while generalization measures reliable performance on new data.
  2. Underfitting means the model has not captured all relevant training patterns.
  3. Overfitting begins when training-specific details improve familiar examples without transferring reliably.
  4. Training and validation performance often improve together at first, but validation performance can later stall and worsen.
  5. More training data is the strongest general remedy; when it is unavailable, limit the information the model can retain.

Key Takeaways

  • Optimization and generalization describe different kinds of model performance.
  • A model can move from underfitting to useful learning and then into overfitting as training continues.
  • Validation performance helps identify when additional training has stopped improving generalization.
  • Overfitting occurs when training-specific patterns improve familiar examples but fail to help on new data.
  • More training data or constraints on retained information can reduce overfitting.