Model Capacity
Overfitting is reduced by making memorization less easy and generalizable representations more useful.
Why Flexibility Can Mislead
A model that fits its training data is not necessarily a model that has learned a useful pattern. A highly flexible model may use its many learnable parameters to remember the answer associated with each training example. That can make training performance look strong while weakening performance on examples outside the training set. Model capacity is the amount of flexibility controlled through the number of learnable parameters. The goal is not maximum capacity. The goal is a balance: enough capacity to learn predictive representations, but not so much that memorizing accidental details becomes easy.
Reducing overfitting does not mean making a model learn less carefully. It means making memorization less attractive than learning representations that retain predictive value beyond the training examples.
Capacity as a Learning Pressure
Think of model capacity as the number of memorization resources available to the network. With excessive capacity, the network has enough flexibility to fit individual training cases, including details that do not represent a useful general pattern. With fewer parameters, the network has less room to assign separate behavior to every case. It is pushed toward a compressed representation that continues to provide predictive information about the targets.
Choosing between two model sizes
A large network fits the training set very closely, but its performance on new examples is weaker. A smaller network performs less impressively on the training set while retaining more useful predictive behavior on new examples. Which model is showing the healthier balance?
Inspect the training behavior: The large network has enough flexibility to fit the training examples closely. That observation alone does not establish that it has learned a general pattern.
Inspect behavior beyond training: The smaller network retains more predictive value on examples outside the training set. This indicates that it is relying less on accidental training details.
Identify the balance: The smaller network is the better-balanced choice in this situation, provided it has not become so small that it fails to learn the useful pattern at all.
The appropriate target is neither the largest model nor the smallest possible model. It is a model with enough parameters to learn the task but not enough flexibility to make memorization the easiest solution.
Weight-Based Constraints
Reducing model size is not the only way to discourage memorization. L1 and L2 regularization leave the model architecture in place but add different weight-based costs to the loss function. The regularizer makes certain learned weight configurations more costly, especially configurations involving large weights. Training therefore has an additional pressure: fitting the data is considered together with the cost associated with the weights.
| Strategy | What changes | How it discourages overfitting |
|---|---|---|
| Reduce model size | The number of learnable parameters | The network has fewer resources available for memorization |
| L1 regularization | A weight-based cost is added to the loss | Weight configurations are evaluated with an additional regularization cost |
| L2 regularization | A different weight-based cost is added to the loss | Large weights become costly while the architecture remains in place |
Dropout State Changes
Dropout applies a different kind of pressure. During training, it randomly zeros some features in a layer's output. The available output information therefore changes as training proceeds and can differ across training examples. A network is discouraged from depending on one fragile combination of features that works only for particular training cases. The source describes this changing noise as a way to break up happenstance patterns that might otherwise be memorized.
What do you think happens?
What should happen to the network's available output features during inference after dropout was used during training?
Reveal answer
Answer: The full network should be used with scaled outputs.
Dropout changes layer outputs by randomly zeroing features during training. Inference uses scaled outputs without dropping units, so prediction does not use a newly dropped subset of features.
Dropout is not simply a permanent deletion of units. Training uses randomly changed layer outputs, whereas inference uses the full network with scaled outputs.
Selecting a Reduction Strategy
Choose the strategy by identifying which part of learning should receive pressure. If the network has excessive flexibility, reducing the number of learnable parameters directly limits its memorization resources. If the architecture is useful but large learned weights are making accidental relationships attractive, L1 or L2 regularization adds a weight-based cost to the loss while leaving the architecture in place. If the network is depending too heavily on a fragile collection of output features, dropout changes which features are available during training and discourages that dependence.
Common Reasoning Errors
Assuming that the model with the best training fit must be the best model.
Fitting training data is easier than generalizing beyond it. Excessive capacity can make memorization of accidental details easy.
Fix:
Judge capacity by the balance between learning useful predictive representations and avoiding training-set memorization.Treating overfitting reduction as simply making the network learn less.
The target is a balance between overfitting and underfitting, not the smallest possible model.
Fix:
Reduce memorization resources while preserving enough capacity to learn predictive information.Describing L1 and L2 regularization as architecture changes.
Regularization changes the loss through a weight-based cost; it does not reduce the number of parameters.
Fix:
Separate the question of how many parameters exist from the question of how costly certain weights are.Assuming dropout removes the same features permanently during training and inference.
Dropout changes layer outputs during training and uses the full network differently during inference.
Fix:
Keep the training state and inference state distinct when reasoning about dropout.
Practice the Choice
A neural network is overfitting. You want to keep its architecture but make large learned weights contribute an additional cost to the training objective. Which strategy best matches that goal, and how is it different from reducing model size?
Hints
- Look for the strategy that changes the loss through a weight-based cost.
- Compare that with the strategy that changes the number of learnable parameters.
A network seems to depend on a fragile combination of features that works for particular training cases. Which strategy changes the available output features during training, and what happens to the outputs during inference?
Hints
- Identify the method that randomly zeros features during training.
- Inference uses the full network with scaled outputs.
Key Takeaways
- Model capacity is controlled through the number of learnable parameters, and excessive capacity can make memorization easier than generalization.
- Reducing model size should aim for a balance: enough capacity to learn predictive representations, but not so much that the training set is memorized.
- L1 and L2 regularization keep the architecture in place while adding different weight-based costs to the loss function.
- Dropout randomly zeros features during training and uses scaled outputs from the full network during inference.
- Choose among model-size reduction, weight regularization, and dropout by identifying whether the main pressure should act on parameter count, weight cost, or dependence on particular output features.
Key Takeaways
- Excessive model capacity can let a network memorize training examples instead of learning patterns that generalize.
- Reducing model size creates a balance between overfitting and underfitting rather than simply making the model learn less.
- L1 and L2 regularization add different weight-based costs to the loss while leaving the architecture in place.
- Dropout changes layer outputs by randomly zeroing features during training; inference uses scaled outputs without dropping units.
- The best overfitting-reduction strategy depends on whether the desired pressure should affect parameter count, weight cost, or feature dependence.