Dropout
Overfitting is reduced by making memorization less easy and generalizable representations more useful.
When Capacity Becomes Memorization
A neural network can fit its training data without learning patterns that remain useful on new examples. This happens when the model has enough flexibility to memorize training examples rather than form generalizable representations. Reducing overfitting means making memorization less easy and making useful, generalizable representations more valuable.
Model capacity is controlled through the number of learnable parameters. Too little capacity can lead to underfitting, while excessive capacity can make memorization easier. The target is a balance: enough capacity to learn predictive structure, but not so much flexibility that accidental details of the training set become the model's main strategy.
Choosing between two model sizes
A model is flexible enough to fit its training examples very closely, but its useful behavior on new examples is poor. Consider whether reducing model size would help.
Identify the pressure: The model has enough capacity to memorize training examples. Fitting the training data is therefore easier than learning patterns that generalize.
Change the capacity: Reducing the number of learnable parameters removes some memorization resources. This can make accidental details less useful to the model.
Check the opposite risk: Reducing capacity too far can cause underfitting. The objective is not the smallest possible model; it is a balance between overfitting and underfitting.
Reducing model size is an overfitting-reduction strategy when the original model has excessive capacity, but the reduction must preserve enough capacity for useful representations.
Three Pressures Against Overfitting
Overfitting can be addressed by changing different parts of the learning problem. Reducing model size changes how many parameters exist. Weight regularization keeps the architecture in place but adds a cost based on the weights. Dropout changes which output features are available during training. These methods share a goal, but they apply different pressures.
| Strategy | What changes | Pressure created |
|---|---|---|
| Reduce model size | Number of learnable parameters | Fewer resources for memorization |
| L1 regularization | Loss receives a weight-based cost | Weight values are made costly through the L1 regularizer |
| L2 regularization | Loss receives a different weight-based cost | Weight values are made costly through the L2 regularizer |
| Dropout | Some layer-output features during training | Fragile feature combinations become less dependable |
L1 and L2 regularization both add weight-based costs to the loss function, but they are different regularizers. The important point is that the loss no longer reflects only how well the model fits its targets; it also includes a cost associated with the layer's weights. In the source example, the L2 regularizer is attached to a layer's kernel weights, so that layer contributes a cost based on those weights.
Dropout During Training
Dropout changes a layer's outputs during training by randomly zeroing features. Because a different subset of features can be removed for different training examples, the network is discouraged from depending on a fragile combination that works only for particular training cases. In this way, dropout makes happenstance patterns harder to memorize.
The training and inference behaviors are different. During training, dropout randomly zeroes features in the layer output. During inference, units are not dropped; instead, scaled outputs are used. Therefore, dropout is not a permanent removal of parts of the trained architecture. It is a training-time change to the information available to the network, followed by a different inference-time use of the network.
Why changing available features matters
Suppose a network can rely on one fragile combination of output features to recognize a training example. What pressure does dropout add during training?
Remove part of the output: Dropout randomly zeroes some output features during training, so the fragile combination is not always available.
Make alternate information useful: The network is discouraged from depending too heavily on that single combination and is pushed toward representations that retain predictive value when some features are unavailable.
Change behavior at inference: At inference, units are not dropped. Scaled outputs are used instead, so the full network is used differently from the randomly altered training output.
Dropout reduces the attractiveness of memorizing accidental feature combinations by changing the available output information during training.
Recurrent Layers and Capacity
Recurrent dropout is dropout used in recurrent layers to fight overfitting. It should be understood as a dropout operation associated specifically with a recurrent layer, not as a method for adding representational depth.
Stacking recurrent layers addresses a different design goal. It means using recurrent layers at multiple levels instead of relying on only one recurrent layer. The result is greater representational power, but the source also emphasizes the cost: higher computational loads.
Choosing the Technique
Choose the technique by starting with the design goal. If the main problem is overfitting in a recurrent layer, recurrent dropout is the directly relevant concept. If the main goal is to increase the network's representational power, stacking recurrent layers addresses that goal, while also increasing computational load.
| Question | Better match | Reason |
|---|---|---|
| Is the recurrent network memorizing training information too easily? | Recurrent dropout | It applies dropout in recurrent layers to fight overfitting. |
| Does the network need greater representational power? | Stacked recurrent layers | Stacking increases representational power. |
| Can the system accept additional computational work? | Stacked recurrent layers, if greater capacity is the goal | Stacking also increases computational loads. |
The two techniques should not be selected as if they were identical replacements.
A recurrent network is already sufficiently expressive, but it is overfitting its training data. Would recurrent dropout or stacking recurrent layers be the closer match to the stated goal? Explain what changes and what trade-off the choice introduces.
Hints
- Start with the primary goal rather than the name of the architecture.
- Separate reducing overfitting from increasing representational power.
- Recall that stacking also increases computational loads.
Common Misunderstandings
Assuming that better training fit always means better generalization.
Fitting training data is easier than generalizing beyond it, and excessive capacity can support memorization.
Fix:
Evaluate the balance between overfitting and underfitting rather than treating training fit as the only objective.Treating dropout as permanent removal of units.
Dropout randomly zeroes features during training, while inference uses scaled outputs without dropping units.
Fix:
Keep training behavior and inference behavior separate.Treating L1 and L2 regularization as the same cost.
The source states that L1 and L2 add different weight-based costs to the loss function.
Fix:
Describe both as loss modifications based on weights, while preserving the distinction between the two regularizers.Using stacking as if its primary purpose were to fight overfitting.
Stacking recurrent layers increases representational power and computational loads; recurrent dropout is the technique described as fighting overfitting in recurrent layers.
Fix:
Choose stacking for increased representational power and recurrent dropout for the stated overfitting-reduction goal.Assuming that reducing model size can never cause a problem.
The target is a balance between overfitting and underfitting.
Fix:
Reduce excessive capacity without assuming that the smallest model is automatically best.
Key Takeaways
- Excessive model capacity can make memorization of training examples easier than learning patterns that generalize.
- Reducing model size, adding weight-based regularization, and using dropout create different pressures against overfitting.
- L1 and L2 regularization add different weight-based costs to the loss function.
- Dropout randomly zeroes features during training, while inference uses scaled outputs without dropping units.
- Recurrent dropout is dropout associated with recurrent layers to fight overfitting; stacking recurrent layers increases representational power but also increases computational loads.
Key Takeaways
- Overfitting occurs when a model uses its flexibility to memorize training examples instead of learning generalizable patterns.
- Model size, weight regularization, and dropout reduce overfitting through different mechanisms.
- Dropout changes layer outputs during training but uses scaled outputs without dropping units during inference.
- Recurrent dropout targets overfitting in recurrent layers, whereas stacking recurrent layers targets greater representational power at a higher computational cost.