Machine learning model capacity
Optimization and generalization measure different kinds of model performance.
Two Kinds of Improvement
A model can become better at handling the examples it has already seen while becoming worse at handling new examples. This is possible because optimization and generalization measure different kinds of performance. Optimization concerns improvement on the training data. Generalization concerns how reliably the learned behavior works on data the model has not seen.
The Training Journey
Training commonly begins with agreement between the two measures. As the model starts learning, its performance on training examples improves, and its performance on unseen examples often improves too. During this period, the model may still be underfit because it has not yet captured all the relevant patterns in the training data.
What do you think happens?
After validation performance reaches its strongest point, what can happen if training continues?
Reveal answer
Answer: Training performance can improve while validation performance stalls or worsens
The model may continue improving on familiar training examples while becoming increasingly tailored to details that do not transfer reliably to new data.
Underfitting and Overfitting
Underfitting means that the model has not represented all the relevant patterns in its training data. Further useful progress is still available.
Overfitting begins when the model learns details that belong specifically to the training data but do not provide reliable guidance for new data. Those details may be misleading or irrelevant outside the training set.
These terms describe different stages of the relationship between learning and transfer. Underfitting does not simply mean that the model performs poorly for an unexplained reason. It means the model has not yet captured all the useful structure available in the training data. Overfitting is different: the model has started using training-specific information that does not help it perform reliably on new data.
Useful Patterns and Training Details
A useful pattern captures a relationship that remains helpful when the model encounters new data. A training-specific pattern helps with the examples used during training but does not transfer reliably. The important test is not whether the model can improve its results on familiar examples; it is whether the improvement remains visible on unseen examples.
A topic-classification model
A model is trained to classify the topics of documents. Early in training, both its training performance and its validation performance improve. Later, training performance continues to improve, but validation metrics stop improving and then worsen. What is happening?
Early training: The model is learning patterns that help with the training examples and also appear useful on unseen examples. This is consistent with progress away from underfitting.
Validation peak: Validation performance reaches its strongest point. At this point, the model is showing its best observed generalization during the training process.
Continued training: Training performance keeps improving, but validation performance stalls and then declines. The model is increasingly learning details specific to the training data.
Diagnosis: Optimization is continuing, but generalization is no longer improving. The model has moved into overfitting.
The model should be evaluated by its validation behavior, not by training performance alone. Continued training after the validation peak is associated with overfitting in this example.
Model Capacity and Fit
Model capacity concerns how much information a model can retain or represent. The source describes reducing overfitting partly by limiting what the model can retain or by placing constraints on the information it is allowed to store. When unrestricted memorization is made less available, optimization is pushed toward the most prominent patterns, which have a better chance of generalizing.
Reducing Overfitting
The strongest general remedy described for overfitting is obtaining more training data. More examples give the model more information from which to learn and make it more likely to develop patterns that generalize.
When more data is not possible, reduce the amount of information the model can retain or place constraints on what it is allowed to store. These approaches make unrestricted memorization less available and encourage the optimization process to focus on the most prominent patterns.
- Prefer obtaining more training data when that is possible.
- If more data is unavailable, limit what the model can retain or constrain the information it is allowed to store.
- Track validation performance while training continues.
- Treat a stall or decline in validation performance, alongside continued training improvement, as evidence that the model is moving into overfitting.
- Favor the point in training where validation performance is strongest rather than relying only on the best training performance.
Mistakes to Avoid
Assuming that better training performance always means better model performance.
Optimization and generalization measure different kinds of performance. Improvement on familiar examples can coexist with worsening performance on unseen examples.
Fix:
Compare training behavior with validation behavior.Calling every poorly performing model overfit.
This describes underfitting, not overfitting. Further useful progress is still available.
Fix:
Use underfitting when relevant training patterns have not yet been represented.Treating validation performance as unnecessary once training performance is high.
Validation performance helps reveal when continued training has stopped improving generalization.
Fix:
Continue checking performance on data the model has not seen.Assuming that a training-specific detail is useful because it improves familiar examples.
Overfitting begins when such details replace patterns that transfer reliably.
Fix:
Ask whether the pattern continues to help on unseen data.
Check Your Understanding
A model's training performance improves throughout training. Its validation performance improves at first, reaches a strongest point, and then begins to decline. Identify the training phase before the strongest validation point, the phase after it, and the difference between optimization and generalization in this situation.
Hints
- Before the strongest validation point, the model is still learning useful patterns.
- After the strongest validation point, compare the directions of training and validation performance.
- Optimization concerns training examples; generalization concerns unseen examples.
Practice answer
Interpret the changing training and validation performance.
Before the validation peak: The model is in an underfitting phase if it has not yet captured all relevant patterns and both training and unseen-data performance are improving.
After the validation peak: The model is moving into overfitting when training performance continues to improve but validation performance stalls and then worsens.
Meaning of the two measures: Optimization is continuing because performance on training examples improves. Generalization is no longer improving because performance on unseen examples has stopped improving or declined.
The model has changed from underfitting to overfitting as training continued.
Key Takeaways
- Optimization measures how performance changes on training examples, while generalization concerns performance on unseen examples.
- Underfitting means that relevant training patterns have not yet been captured.
- Overfitting begins when training-specific details improve familiar examples but fail to transfer reliably to new data.
- A model can move from underfitting to overfitting: both measures may improve at first, then validation performance can stall or worsen while training performance continues improving.
- More training data is the strongest general remedy described; when it is unavailable, limiting what the model can retain or constraining stored information can reduce unrestricted memorization.
Key Takeaways
- Optimization and generalization measure different kinds of model performance.
- Underfitting can become overfitting as training continues beyond the point of strongest validation performance.
- Useful patterns transfer to new data, while training-specific details do not transfer reliably.
- More training data, reduced information retention, and constraints on stored information are practical ways to reduce overfitting.