Concepts / Dropout

Dropout

Overfitting is reduced by making memorization less easy and generalizable representations more useful.

  • Programming

When Capacity Becomes Memorization

A neural network can fit its training data without learning patterns that remain useful on new examples. This happens when the model has enough flexibility to memorize training examples rather than form generalizable representations. Reducing overfitting means making memorization less easy and making useful, generalizable representations more valuable.

too little capacitysuitable capacityexcess flexibilitySmaller modelfewer parametersUnderfittinginsufficient flexibilityBalanced capacityuseful representationsGeneralizationpatterns beyond trainingexamplesLarger modelmore parametersMemorizationtraining examples
How does changing model size alter the balance between underfitting, memorization, and generalization?

Model capacity is controlled through the number of learnable parameters. Too little capacity can lead to underfitting, while excessive capacity can make memorization easier. The target is a balance: enough capacity to learn predictive structure, but not so much flexibility that accidental details of the training set become the model's main strategy.

Choosing between two model sizes

A model is flexible enough to fit its training examples very closely, but its useful behavior on new examples is poor. Consider whether reducing model size would help.

Identify the pressure: The model has enough capacity to memorize training examples. Fitting the training data is therefore easier than learning patterns that generalize.

Change the capacity: Reducing the number of learnable parameters removes some memorization resources. This can make accidental details less useful to the model.

Check the opposite risk: Reducing capacity too far can cause underfitting. The objective is not the smallest possible model; it is a balance between overfitting and underfitting.

Reducing model size is an overfitting-reduction strategy when the original model has excessive capacity, but the reduction must preserve enough capacity for useful representations.

Three Pressures Against Overfitting

Overfitting can be addressed by changing different parts of the learning problem. Reducing model size changes how many parameters exist. Weight regularization keeps the architecture in place but adds a cost based on the weights. Dropout changes which output features are available during training. These methods share a goal, but they apply different pressures.

StrategyWhat changesPressure created
Reduce model sizeNumber of learnable parametersFewer resources for memorization
L1 regularizationLoss receives a weight-based costWeight values are made costly through the L1 regularizer
L2 regularizationLoss receives a different weight-based costWeight values are made costly through the L2 regularizer
DropoutSome layer-output features during trainingFragile feature combinations become less dependable

L1 and L2 regularization both add weight-based costs to the loss function, but they are different regularizers. The important point is that the loss no longer reflects only how well the model fits its targets; it also includes a cost associated with the layer's weights. In the source example, the L2 regularizer is attached to a layer's kernel weights, so that layer contributes a cost based on those weights.

combined withaddscombined withaddsData-fit lossfit to targetsL1 costweight-based costLoss with L1fit cost plus L1 costL2 costdifferent weight-based costLoss with L2fit cost plus L2 cost
What do L1 and L2 regularization change in the learning objective?

Dropout During Training

Dropout changes a layer's outputs during training by randomly zeroing features. Because a different subset of features can be removed for different training examples, the network is discouraged from depending on a fragile combination that works only for particular training cases. In this way, dropout makes happenstance patterns harder to memorize.

randomly zeroes featuresusesTrainingfeatures randomly zeroedPartial outputsome features unavailableInferenceunits are not droppedScaled outputfull network used
What changes when units are randomly dropped during training, and how is the network used differently during inference?

The training and inference behaviors are different. During training, dropout randomly zeroes features in the layer output. During inference, units are not dropped; instead, scaled outputs are used. Therefore, dropout is not a permanent removal of parts of the trained architecture. It is a training-time change to the information available to the network, followed by a different inference-time use of the network.

Why changing available features matters

Suppose a network can rely on one fragile combination of output features to recognize a training example. What pressure does dropout add during training?

Remove part of the output: Dropout randomly zeroes some output features during training, so the fragile combination is not always available.

Make alternate information useful: The network is discouraged from depending too heavily on that single combination and is pushed toward representations that retain predictive value when some features are unavailable.

Change behavior at inference: At inference, units are not dropped. Scaled outputs are used instead, so the full network is used differently from the randomly altered training output.

Dropout reduces the attractiveness of memorizing accidental feature combinations by changing the available output information during training.

Recurrent Layers and Capacity

Recurrent dropout is dropout used in recurrent layers to fight overfitting. It should be understood as a dropout operation associated specifically with a recurrent layer, not as a method for adding representational depth.

processedproducesprocessed at a later timestepproducesInput at t1sequence informationRecurrent layerrecurrent dropoutOutput at t1training output affectedInput at t2later sequence informationOutput at t2training output affected
Where does the overfitting-reduction idea apply as information moves through a recurrent layer over timesteps?

Stacking recurrent layers addresses a different design goal. It means using recurrent layers at multiple levels instead of relying on only one recurrent layer. The result is greater representational power, but the source also emphasizes the cost: higher computational loads.

passes throughpasses to next levelsupportsadds toSequence inputinput informationRecurrent layer 1first recurrent levelRecurrent layer 2additional recurrent levelRepresentationalpowerincreasesComputational loadincreases
How does information move through multiple recurrent layers, and why does stacking increase both representational power and computational load?

Choosing the Technique

Choose the technique by starting with the design goal. If the main problem is overfitting in a recurrent layer, recurrent dropout is the directly relevant concept. If the main goal is to increase the network's representational power, stacking recurrent layers addresses that goal, while also increasing computational load.

if the goal ischooseif the goal ischoosePrimary goalwhat problem needs solving?Reduce overfittingrecurrent layerRecurrent dropoutdropout associated withrecurrent layerIncreaserepresentationalpowerrecurrent networkStack recurrentlayershigher computational load
How should you choose between reducing overfitting and increasing representational power?
QuestionBetter matchReason
Is the recurrent network memorizing training information too easily?Recurrent dropoutIt applies dropout in recurrent layers to fight overfitting.
Does the network need greater representational power?Stacked recurrent layersStacking increases representational power.
Can the system accept additional computational work?Stacked recurrent layers, if greater capacity is the goalStacking also increases computational loads.

The two techniques should not be selected as if they were identical replacements.

MEDIUM

A recurrent network is already sufficiently expressive, but it is overfitting its training data. Would recurrent dropout or stacking recurrent layers be the closer match to the stated goal? Explain what changes and what trade-off the choice introduces.

Hints
  • Start with the primary goal rather than the name of the architecture.
  • Separate reducing overfitting from increasing representational power.
  • Recall that stacking also increases computational loads.

Common Misunderstandings

  • Assuming that better training fit always means better generalization.

    Fitting training data is easier than generalizing beyond it, and excessive capacity can support memorization.

    Fix: Evaluate the balance between overfitting and underfitting rather than treating training fit as the only objective.

  • Treating dropout as permanent removal of units.

    Dropout randomly zeroes features during training, while inference uses scaled outputs without dropping units.

    Fix: Keep training behavior and inference behavior separate.

  • Treating L1 and L2 regularization as the same cost.

    The source states that L1 and L2 add different weight-based costs to the loss function.

    Fix: Describe both as loss modifications based on weights, while preserving the distinction between the two regularizers.

  • Using stacking as if its primary purpose were to fight overfitting.

    Stacking recurrent layers increases representational power and computational loads; recurrent dropout is the technique described as fighting overfitting in recurrent layers.

    Fix: Choose stacking for increased representational power and recurrent dropout for the stated overfitting-reduction goal.

  • Assuming that reducing model size can never cause a problem.

    The target is a balance between overfitting and underfitting.

    Fix: Reduce excessive capacity without assuming that the smallest model is automatically best.

Key Takeaways

  1. Excessive model capacity can make memorization of training examples easier than learning patterns that generalize.
  2. Reducing model size, adding weight-based regularization, and using dropout create different pressures against overfitting.
  3. L1 and L2 regularization add different weight-based costs to the loss function.
  4. Dropout randomly zeroes features during training, while inference uses scaled outputs without dropping units.
  5. Recurrent dropout is dropout associated with recurrent layers to fight overfitting; stacking recurrent layers increases representational power but also increases computational loads.

Key Takeaways

  • Overfitting occurs when a model uses its flexibility to memorize training examples instead of learning generalizable patterns.
  • Model size, weight regularization, and dropout reduce overfitting through different mechanisms.
  • Dropout changes layer outputs during training but uses scaled outputs without dropping units during inference.
  • Recurrent dropout targets overfitting in recurrent layers, whereas stacking recurrent layers targets greater representational power at a higher computational cost.