Concepts / Model Capacity

Model Capacity

Overfitting is reduced by making memorization less easy and generalizable representations more useful.

  • Programming

Why Flexibility Can Mislead

A model that fits its training data is not necessarily a model that has learned a useful pattern. A highly flexible model may use its many learnable parameters to remember the answer associated with each training example. That can make training performance look strong while weakening performance on examples outside the training set. Model capacity is the amount of flexibility controlled through the number of learnable parameters. The goal is not maximum capacity. The goal is a balance: enough capacity to learn predictive representations, but not so much that memorizing accidental details becomes easy.

Reducing overfitting does not mean making a model learn less carefully. It means making memorization less attractive than learning representations that retain predictive value beyond the training examples.

fewer memorization resourcesmore flexibilityFewer parameterscompressed representationsGeneralizable patternpredictive informationMany parametersgreater flexibilityMemorized examplesaccidental details
How does increasing model capacity change what a model can memorize, and why can a smaller model generalize better?

Capacity as a Learning Pressure

Think of model capacity as the number of memorization resources available to the network. With excessive capacity, the network has enough flexibility to fit individual training cases, including details that do not represent a useful general pattern. With fewer parameters, the network has less room to assign separate behavior to every case. It is pushed toward a compressed representation that continues to provide predictive information about the targets.

Choosing between two model sizes

A large network fits the training set very closely, but its performance on new examples is weaker. A smaller network performs less impressively on the training set while retaining more useful predictive behavior on new examples. Which model is showing the healthier balance?

Inspect the training behavior: The large network has enough flexibility to fit the training examples closely. That observation alone does not establish that it has learned a general pattern.

Inspect behavior beyond training: The smaller network retains more predictive value on examples outside the training set. This indicates that it is relying less on accidental training details.

Identify the balance: The smaller network is the better-balanced choice in this situation, provided it has not become so small that it fails to learn the useful pattern at all.

The appropriate target is neither the largest model nor the smallest possible model. It is a model with enough parameters to learn the task but not enough flexibility to make memorization the easiest solution.

reduce model sizereduce furtherExcessive capacitymemorization is easyBalanced capacitypattern learning is favoredInsufficient capacityuseful pattern is hard tolearn
What changes as model size decreases from excessive capacity toward a balanced size and then toward insufficient capacity?

Weight-Based Constraints

Reducing model size is not the only way to discourage memorization. L1 and L2 regularization leave the model architecture in place but add different weight-based costs to the loss function. The regularizer makes certain learned weight configurations more costly, especially configurations involving large weights. Training therefore has an additional pressure: fitting the data is considered together with the cost associated with the weights.

StrategyWhat changesHow it discourages overfitting
Reduce model sizeThe number of learnable parametersThe network has fewer resources available for memorization
L1 regularizationA weight-based cost is added to the lossWeight configurations are evaluated with an additional regularization cost
L2 regularizationA different weight-based cost is added to the lossLarge weights become costly while the architecture remains in place
adds costadds costL1 regularizationweight-based costLoss with L1 costdata loss plus added costL2 regularizationdifferent weight-based costLoss with L2 costdata loss plus added cost
How do L1 and L2 regularization change the training objective without changing the model architecture?

Dropout State Changes

Dropout applies a different kind of pressure. During training, it randomly zeros some features in a layer's output. The available output information therefore changes as training proceeds and can differ across training examples. A network is discouraged from depending on one fragile combination of features that works only for particular training cases. The source describes this changing noise as a way to break up happenstance patterns that might otherwise be memorized.

What do you think happens?

What should happen to the network's available output features during inference after dropout was used during training?

  • A new random subset of features should be removed
  • The full network should be used with scaled outputs
  • All learned weights should be discarded
Reveal answer

Answer: The full network should be used with scaled outputs.

Dropout changes layer outputs by randomly zeroing features during training. Inference uses scaled outputs without dropping units, so prediction does not use a newly dropped subset of features.

traininginferenceLayer outputsfeatures availableRandomly zeroedfeaturestraining stateLayer outputsfull networkScaled outputsinference state
What changes in the network when dropout is applied during training, and how is the full network used differently during inference?

Dropout is not simply a permanent deletion of units. Training uses randomly changed layer outputs, whereas inference uses the full network with scaled outputs.

Selecting a Reduction Strategy

Choose the strategy by identifying which part of learning should receive pressure. If the network has excessive flexibility, reducing the number of learnable parameters directly limits its memorization resources. If the architecture is useful but large learned weights are making accidental relationships attractive, L1 or L2 regularization adds a weight-based cost to the loss while leaving the architecture in place. If the network is depending too heavily on a fragile collection of output features, dropout changes which features are available during training and discourages that dependence.

inspect parametersyesnoyesnoyesOverfittingpressurememorization is too easyModel sizetoo many parametersReduce model sizefewer parametersWeight behaviorlarge weights are costlyL1 or L2regularizationadded weight-based costFeature dependencefragile feature combinationDropoutchanging training outputs
How should a practitioner choose among reducing model size, applying weight regularization, and using dropout?

Common Reasoning Errors

  • Assuming that the model with the best training fit must be the best model.

    Fitting training data is easier than generalizing beyond it. Excessive capacity can make memorization of accidental details easy.

    Fix: Judge capacity by the balance between learning useful predictive representations and avoiding training-set memorization.

  • Treating overfitting reduction as simply making the network learn less.

    The target is a balance between overfitting and underfitting, not the smallest possible model.

    Fix: Reduce memorization resources while preserving enough capacity to learn predictive information.

  • Describing L1 and L2 regularization as architecture changes.

    Regularization changes the loss through a weight-based cost; it does not reduce the number of parameters.

    Fix: Separate the question of how many parameters exist from the question of how costly certain weights are.

  • Assuming dropout removes the same features permanently during training and inference.

    Dropout changes layer outputs during training and uses the full network differently during inference.

    Fix: Keep the training state and inference state distinct when reasoning about dropout.

Practice the Choice

MEDIUM

A neural network is overfitting. You want to keep its architecture but make large learned weights contribute an additional cost to the training objective. Which strategy best matches that goal, and how is it different from reducing model size?

Hints
  • Look for the strategy that changes the loss through a weight-based cost.
  • Compare that with the strategy that changes the number of learnable parameters.
MEDIUM

A network seems to depend on a fragile combination of features that works for particular training cases. Which strategy changes the available output features during training, and what happens to the outputs during inference?

Hints
  • Identify the method that randomly zeros features during training.
  • Inference uses the full network with scaled outputs.

Key Takeaways

  1. Model capacity is controlled through the number of learnable parameters, and excessive capacity can make memorization easier than generalization.
  2. Reducing model size should aim for a balance: enough capacity to learn predictive representations, but not so much that the training set is memorized.
  3. L1 and L2 regularization keep the architecture in place while adding different weight-based costs to the loss function.
  4. Dropout randomly zeros features during training and uses scaled outputs from the full network during inference.
  5. Choose among model-size reduction, weight regularization, and dropout by identifying whether the main pressure should act on parameter count, weight cost, or dependence on particular output features.

Key Takeaways

  • Excessive model capacity can let a network memorize training examples instead of learning patterns that generalize.
  • Reducing model size creates a balance between overfitting and underfitting rather than simply making the model learn less.
  • L1 and L2 regularization add different weight-based costs to the loss while leaving the architecture in place.
  • Dropout changes layer outputs by randomly zeroing features during training; inference uses scaled outputs without dropping units.
  • The best overfitting-reduction strategy depends on whether the desired pressure should affect parameter count, weight cost, or feature dependence.