Overfitting and Underfitting
Overfitting is reduced by making memorization less easy and generalizable representations more useful.
The Training-Set Trap
A model can become better at the examples it has already seen while becoming less dependable on new examples. This happens because fitting familiar data is easier than learning patterns that generalize beyond it. The central challenge is not to make the model learn less carefully. It is to prevent the model from using its flexibility to memorize training-specific details.
Overfitting is reduced by making memorization less easy and making generalizable representations more useful.
The same issue can appear in different tasks, including movie-review prediction, topic classification, and house-price regression. In each case, a model may learn details specific to its training examples that do not provide reliable guidance for new data.
Capacity and Memorization
Model capacity is controlled through the number of learnable parameters. A model with too little capacity may fail to capture relevant patterns and is underfit. As capacity increases, the model can represent more useful structure. If it has too much flexibility, it can also use that flexibility to memorize training examples and training-specific details. The target is a balance between underfitting and overfitting.
The visual shows why increasing capacity is not automatically beneficial. Too little capacity leaves relevant patterns unrepresented. Balanced capacity supports useful representations. Excessive capacity can make unrestricted memorization easier.
Three Training Checkpoints
Following a Model Through Training
A model is observed at three points during training. Determine whether it is underfitting, generalizing usefully, or moving into overfitting.
Beginning: The model has not yet captured all the relevant patterns in its training data. This is the underfitting phase, because useful progress is still available.
Middle: Training performance and performance on unseen data both improve. The model is learning patterns that help beyond the examples used for training.
Later: Training performance continues to improve, but validation performance stalls and then worsens. The model is learning details specific to the training data.
The model moves from underfitting toward useful generalization and then into overfitting.
This progression is important because overfitting is not always visible from training performance alone. At first, training and validation behavior often improve together. Later, optimization can continue while generalization stops improving.
Optimization Versus Generalization
| Measure | What it describes | What continued improvement can mean |
|---|---|---|
| Optimization | Performance on the data used for training | The model is fitting the training examples more closely |
| Generalization | Performance on data the model has not seen | The model is becoming more dependable beyond the training set |
Optimization and generalization measure different kinds of model performance. Reducing training loss is evidence that optimization is continuing, but it does not prove that the model is learning more useful behavior for unseen data. Validation performance helps reveal whether continued training is still improving generalization.
Pressure Against Memorization
When more data cannot be obtained, overfitting can be addressed by changing how much information the model can store or by placing constraints on the information it is allowed to store. The main controls in this lesson are model size, weight regularization, and dropout. They work differently, but all make accidental memorization less attractive than learning representations that retain predictive value.
| Strategy | What changes | Shared purpose |
|---|---|---|
| Reduce model size | The number of learnable parameters | Limit memorization resources |
| Weight regularization | The loss receives an additional weight-based cost | Make certain weight choices less attractive |
| Dropout | Some output features are removed during training | Discourage dependence on fragile feature combinations |
| Obtain more training data | The model learns from more examples | Increase the chance of learning generalizable patterns |
These strategies reduce overfitting through different mechanisms.
Weight-Based Costs
L1 and L2 regularization add different weight-based costs to the loss function. The regularizer is attached to a layer's kernel weights, so the layer contributes an additional cost based on those weights. This leaves the model architecture in place while making certain weight choices less attractive during learning.
Dropout Changes by Phase
Dropout changes layer outputs during training by randomly zeroing features. Because a different subset of features can be removed for different training examples, the network is discouraged from depending on a fragile combination of features that only works for particular training cases. This breaks up happenstance patterns that the network might otherwise memorize.
Inference uses the full network differently: units are not dropped, and outputs are scaled rather than being produced by randomly removing features. Therefore, training and inference do not present exactly the same output behavior to the model.
Selecting a Remedy
Start by comparing training and validation behavior. If more training data is possible, it is the strongest general remedy because additional examples make generalizable patterns more likely. If more data is not available, reduce memorization resources or constrain what the model can store. Reducing model size changes the number of parameters. Regularization keeps the architecture but adds a weight-based cost. Dropout changes which output features are available during training.
Common Misreadings
Treating lower training loss as proof that generalization is improving.
Optimization and generalization measure different kinds of performance.
Fix:
Inspect validation performance alongside training performance.Calling every poorly performing model overfit.
That describes underfitting, not overfitting.
Fix:
Ask whether useful progress is still missing or whether training-specific details have begun to dominate.Assuming that reducing overfitting means making the model learn less carefully.
The goal is to prevent memorization and encourage useful representations, not to abandon careful learning.
Fix:
Describe the intervention as changing the pressure on learning.Treating dropout as a permanent reduction in model size.
Dropout changes output information during training rather than permanently removing parameters.
Fix:
Distinguish training-time feature removal from inference-time scaled full-network outputs.Treating L1 and L2 regularization as the same mechanism.
The source distinguishes them as different costs added to the loss function.
Fix:
Remember that both constrain weights through the loss, while L1 and L2 are not identical penalties.
Check Your Reasoning
A topic-classification model continues to improve on its training examples. Its validation performance first improves, then stops improving, and finally declines. Explain what is happening and choose one strategy that could make memorization less attractive.
Hints
- Separate optimization from generalization.
- Identify the phase in which validation performance is strongest.
- Choose from more training data, reducing model size, weight regularization, or dropout, and explain the mechanism of your choice.
What do you think happens?
If training performance keeps improving after validation performance has begun to decline, is optimization still continuing?
Reveal answer
Answer: Yes, while generalization is no longer improving
Training performance can continue to improve even after validation performance stalls or worsens. That means optimization is continuing, but the model is moving toward overfitting.
Key Takeaways
- Underfitting means that the model has not captured all relevant patterns in its training data.
- Overfitting begins when the model learns training-specific details that do not transfer reliably to new data.
- Optimization concerns fitting the training data, while generalization concerns performance on unseen data.
- Model size, weight regularization, and dropout reduce overfitting through different mechanisms.
- Validation performance reveals when continued optimization has stopped improving generalization.
Key Takeaways
- A model can improve on training examples while becoming worse on unseen data.
- Underfitting reflects missing useful patterns; overfitting reflects reliance on training-specific patterns.
- Optimization and generalization must be evaluated separately through training and validation behavior.
- Reducing model size, adding L1 or L2 weight costs, and using dropout constrain memorization in different ways.
- More training data is the strongest general remedy when it can be obtained.