Concepts / Overfitting and Underfitting

Overfitting and Underfitting

Overfitting is reduced by making memorization less easy and generalizable representations more useful.

  • Programming

The Training-Set Trap

A model can become better at the examples it has already seen while becoming less dependable on new examples. This happens because fitting familiar data is easier than learning patterns that generalize beyond it. The central challenge is not to make the model learn less carefully. It is to prevent the model from using its flexibility to memorize training-specific details.

Overfitting is reduced by making memorization less easy and making generalizable representations more useful.

The same issue can appear in different tasks, including movie-review prediction, topic classification, and house-price regression. In each case, a model may learn details specific to its training examples that do not provide reliable guidance for new data.

Capacity and Memorization

Model capacity is controlled through the number of learnable parameters. A model with too little capacity may fail to capture relevant patterns and is underfit. As capacity increases, the model can represent more useful structure. If it has too much flexibility, it can also use that flexibility to memorize training examples and training-specific details. The target is a balance between underfitting and overfitting.

too little capacityuseful representationmemorization becomes easierFew parametersRelevant patterns missedUnderfittingBalanced capacityUseful patterns representedGeneralizationMany parametersTraining details memorizedOverfitting
What changes as model capacity increases, and when does the model shift from missing useful patterns to memorizing training-specific patterns?

The visual shows why increasing capacity is not automatically beneficial. Too little capacity leaves relevant patterns unrepresented. Balanced capacity supports useful representations. Excessive capacity can make unrestricted memorization easier.

Three Training Checkpoints

Following a Model Through Training

A model is observed at three points during training. Determine whether it is underfitting, generalizing usefully, or moving into overfitting.

Beginning: The model has not yet captured all the relevant patterns in its training data. This is the underfitting phase, because useful progress is still available.

Middle: Training performance and performance on unseen data both improve. The model is learning patterns that help beyond the examples used for training.

Later: Training performance continues to improve, but validation performance stalls and then worsens. The model is learning details specific to the training data.

The model moves from underfitting toward useful generalization and then into overfitting.

This progression is important because overfitting is not always visible from training performance alone. At first, training and validation behavior often improve together. Later, optimization can continue while generalization stops improving.

training continuescontinued optimizationnot all useful patterns learnedunseen-data performancetraining-specific detailsEarly trainingTraining improves;validation improvesUnderfittingStrongestvalidationGeneralization is strongestGeneralizationLater trainingTraining improves;validation worsensOverfitting
How do training and validation results change across training, and when does continued optimization begin to hurt generalization?

Optimization Versus Generalization

MeasureWhat it describesWhat continued improvement can mean
OptimizationPerformance on the data used for trainingThe model is fitting the training examples more closely
GeneralizationPerformance on data the model has not seenThe model is becoming more dependable beyond the training set

Optimization and generalization measure different kinds of model performance. Reducing training loss is evidence that optimization is continuing, but it does not prove that the model is learning more useful behavior for unseen data. Validation performance helps reveal whether continued training is still improving generalization.

optimizationgeneralizationTraining dataFamiliar examplesTraining performanceContinues improvingUnseen dataNew examplesValidationperformanceStalls or worsens
How can reducing training loss improve optimization while validation performance stops improving or gets worse?

Pressure Against Memorization

When more data cannot be obtained, overfitting can be addressed by changing how much information the model can store or by placing constraints on the information it is allowed to store. The main controls in this lesson are model size, weight regularization, and dropout. They work differently, but all make accidental memorization less attractive than learning representations that retain predictive value.

StrategyWhat changesShared purpose
Reduce model sizeThe number of learnable parametersLimit memorization resources
Weight regularizationThe loss receives an additional weight-based costMake certain weight choices less attractive
DropoutSome output features are removed during trainingDiscourage dependence on fragile feature combinations
Obtain more training dataThe model learns from more examplesIncrease the chance of learning generalizable patterns

These strategies reduce overfitting through different mechanisms.

fewer storage resourcesadditional costchanging output informationModel sizeParameter countLess memorizationMore useful representationsWeightregularizationWeight-based costDropoutRemoved features duringtraining
How do model size, weight regularization, and dropout place different pressures on learning?

Weight-Based Costs

L1 and L2 regularization add different weight-based costs to the loss function. The regularizer is attached to a layer's kernel weights, so the layer contributes an additional cost based on those weights. This leaves the model architecture in place while making certain weight choices less attractive during learning.

add L1 costadd L2 costattached to weightsattached to weightsBase lossFit to training examplesL1-regularized lossBase loss plus L1 costL1 costWeight-based additionL2-regularized lossBase loss plus L2 costL2 costWeight-based addition
How do L1 and L2 penalties change the loss function and affect the role of weight values?

Dropout Changes by Phase

Dropout changes layer outputs during training by randomly zeroing features. Because a different subset of features can be removed for different training examples, the network is discouraged from depending on a fragile combination of features that only works for particular training cases. This breaks up happenstance patterns that the network might otherwise memorize.

Inference uses the full network differently: units are not dropped, and outputs are scaled rather than being produced by randomly removing features. Therefore, training and inference do not present exactly the same output behavior to the model.

randomly zero featuresuse full networkLayer featuresTraining phaseRandom feature subsetSome features zeroedLayer featuresInference phaseScaled outputsNo units dropped
Which units are removed during training, and how is the full network used differently during inference?

Selecting a Remedy

if data can be obtainedlimit storage resourcesconstrain weightsdiscourage feature dependenceObserve validationCompare with trainingMore training dataStrongest general remedyReduce model sizeFewer parametersWeight regularizationAdditional weight costDropoutChanging features duringtraining
Given the model, dataset, and training-validation behavior, which strategy should be considered next?

Start by comparing training and validation behavior. If more training data is possible, it is the strongest general remedy because additional examples make generalizable patterns more likely. If more data is not available, reduce memorization resources or constrain what the model can store. Reducing model size changes the number of parameters. Regularization keeps the architecture but adds a weight-based cost. Dropout changes which output features are available during training.

Common Misreadings

  • Treating lower training loss as proof that generalization is improving.

    Optimization and generalization measure different kinds of performance.

    Fix: Inspect validation performance alongside training performance.

  • Calling every poorly performing model overfit.

    That describes underfitting, not overfitting.

    Fix: Ask whether useful progress is still missing or whether training-specific details have begun to dominate.

  • Assuming that reducing overfitting means making the model learn less carefully.

    The goal is to prevent memorization and encourage useful representations, not to abandon careful learning.

    Fix: Describe the intervention as changing the pressure on learning.

  • Treating dropout as a permanent reduction in model size.

    Dropout changes output information during training rather than permanently removing parameters.

    Fix: Distinguish training-time feature removal from inference-time scaled full-network outputs.

  • Treating L1 and L2 regularization as the same mechanism.

    The source distinguishes them as different costs added to the loss function.

    Fix: Remember that both constrain weights through the loss, while L1 and L2 are not identical penalties.

Check Your Reasoning

MEDIUM

A topic-classification model continues to improve on its training examples. Its validation performance first improves, then stops improving, and finally declines. Explain what is happening and choose one strategy that could make memorization less attractive.

Hints
  • Separate optimization from generalization.
  • Identify the phase in which validation performance is strongest.
  • Choose from more training data, reducing model size, weight regularization, or dropout, and explain the mechanism of your choice.

What do you think happens?

If training performance keeps improving after validation performance has begun to decline, is optimization still continuing?

  • Yes, while generalization is no longer improving
  • No, because validation performance declined
  • Only if the model has fewer parameters
Reveal answer

Answer: Yes, while generalization is no longer improving

Training performance can continue to improve even after validation performance stalls or worsens. That means optimization is continuing, but the model is moving toward overfitting.

Key Takeaways

  1. Underfitting means that the model has not captured all relevant patterns in its training data.
  2. Overfitting begins when the model learns training-specific details that do not transfer reliably to new data.
  3. Optimization concerns fitting the training data, while generalization concerns performance on unseen data.
  4. Model size, weight regularization, and dropout reduce overfitting through different mechanisms.
  5. Validation performance reveals when continued optimization has stopped improving generalization.

Key Takeaways

  • A model can improve on training examples while becoming worse on unseen data.
  • Underfitting reflects missing useful patterns; overfitting reflects reliance on training-specific patterns.
  • Optimization and generalization must be evaluated separately through training and validation behavior.
  • Reducing model size, adding L1 or L2 weight costs, and using dropout constrain memorization in different ways.
  • More training data is the strongest general remedy when it can be obtained.