Weight Regularization
Overfitting is reduced by making memorization less easy and generalizable representations more useful.
When Flexibility Becomes Memorization
A neural network must fit its training data, but fitting the training data is easier than generalizing beyond it. If a model has excessive capacity, it may use its flexibility to remember answers associated with individual training examples rather than learning representations that remain useful for new examples. Weight regularization is part of a broader effort to make memorization less easy and generalizable representations more useful.
The target is not the smallest possible model. The target is a balance: enough capacity to learn useful patterns, but not so much flexibility that memorization becomes easy.
Capacity as a Learning Constraint
Model capacity is controlled through the number of learnable parameters. A model with fewer parameters has fewer resources available for memorization. This does not mean that every smaller model will generalize well: reducing capacity too far can lead to underfitting, where the model lacks enough flexibility to learn useful patterns. The practical goal is to preserve enough capacity for predictive representations while limiting unnecessary flexibility.
Choosing Between Two Model Sizes
A large network appears able to memorize training examples, while a much smaller network may fail to learn useful patterns. What is the reasoning behind trying a reduced model size?
Identify the source of flexibility: The number of learnable parameters controls model capacity. The larger network has more parameters available for fitting details of the training set.
Limit memorization resources: Reducing the number of parameters makes it harder for the model to use excessive flexibility to memorize individual training examples.
Check for useful representation: The reduced model still needs enough capacity to learn predictive patterns. If capacity is reduced too far, the model can move toward underfitting.
Reducing model size is a balancing decision: limit memorization without removing the capacity needed to learn generalizable patterns.
Penalizing Large Weights
Weight regularization leaves the model architecture in place but changes the learning pressure. In addition to the ordinary cost associated with the model's predictions, the loss function receives an additional weight-based cost. This makes large weights less attractive during learning. The regularizer is attached to the layer's kernel weights, so the layer contributes a cost based on those weights.
| Technique | What changes | Learning pressure |
|---|---|---|
| L1 regularization | A distinct weight-based cost is added to the loss | The loss includes an additional cost associated with the layer's weights |
| L2 regularization | A different weight-based cost is added to the loss | The loss includes an additional cost associated with the layer's weights |
Reading the Regularization Choice
A layer is configured with an L2 regularizer and the value 0.001. What structural idea should you identify?
Locate the regularizer: The regularizer is attached to the layer rather than changing the number of layers or parameters.
Identify the affected weights: The regularizer is attached to the layer's kernel weights.
Identify the added pressure: The layer contributes an additional loss cost based on those weights. The value 0.001 is the supplied value for the L2 regularizer in the source example.
The structural effect is an added weight-based cost in the loss; the architecture remains in place.
Dropout as Changing Availability
Dropout works through the information available at a layer's output. During training, it randomly zeros some features, so a particular training example does not always present the network with the same complete set of output features. Because a different subset may be removed for different examples, the network is discouraged from depending on a fragile combination of features that works only for particular training cases.
| Phase | What happens to layer outputs | Purpose or effect |
|---|---|---|
| Training | Some features are randomly zeroed | Discourages dependence on fragile feature combinations |
| Inference | Units are not dropped; outputs are scaled | Uses scaled outputs rather than training-time dropping |
Selecting the Intervention
These strategies share a goal but do not make the same change. Reducing model size changes how many learnable parameters exist. L1 and L2 regularization keep the architecture in place while adding different weight-based costs to the loss. Dropout changes which output features are available during training and uses scaled outputs during inference.
Matching the Strategy to the Problem
A network appears too flexible and is memorizing training examples. Decide what each strategy would change before choosing one.
Choose model-size reduction when capacity is the concern: This directly changes the number of learnable parameters, reducing the model's available memorization resources.
Choose weight regularization when weight size should carry a cost: This keeps the architecture in place while adding a weight-based cost to the loss. L1 and L2 provide different versions of that pressure.
Choose dropout when dependence on feature combinations is the concern: During training, dropout randomly zeros features so the network cannot rely as easily on one fragile combination for every training example.
Check the training and inference distinction: Dropout changes outputs during training, while inference uses scaled outputs without dropping units.
Select the strategy by the mechanism you want to change: parameter count, weight-based loss pressure, or training-time feature availability.
Mistakes Beginners Make
Assuming that more parameters are always better
Excessive capacity can make memorization of training examples easier than learning patterns that generalize.
Fix:
Treat model capacity as a balance: retain enough parameters for useful learning while limiting unnecessary memorization resources.Reducing model size as far as possible
Too little capacity can produce underfitting.
Fix:
Reduce capacity enough to make memorization less easy, but preserve enough flexibility for generalizable representations.Treating L1 and L2 as identical
The source distinguishes them as different weight-based costs added to the loss function.
Fix:
Remember that both act through the loss, but L1 and L2 apply different weight-based costs.Assuming dropout permanently removes units
Dropout randomly zeros features during training, whereas inference uses scaled outputs without dropping units.
Fix:
Separate the training behavior from the inference behavior.Thinking regularization changes the architecture
Weight regularization leaves the architecture in place and changes the loss through an added cost based on weights.
Fix:
Distinguish parameter-count changes from weight-cost changes.
Check Your Reasoning
A neural network is memorizing individual training examples. You are allowed to make only one conceptual change. Decide whether each proposal changes model capacity, weight-based loss pressure, or training-time feature availability: reduce the number of parameters, attach an L1 regularizer, attach an L2 regularizer, or add dropout. Then explain why the chosen change could make memorization less attractive.
Hints
- Start by asking what the proposal changes directly.
- Model-size reduction changes the number of learnable parameters.
- L1 and L2 add different weight-based costs to the loss.
- Dropout randomly zeros features during training but uses scaled outputs during inference.
What do you think happens?
A layer uses dropout during training. What should you expect during inference?
Reveal answer
Answer: No units are dropped, and scaled outputs are used
The source distinguishes training-time random feature removal from inference, where dropout does not drop units and scaled outputs are used.
Key Takeaways
- Excessive model capacity can make memorization of training examples easier than learning patterns that generalize.
- Reducing the number of learnable parameters limits memorization resources, but reducing capacity too far can cause underfitting.
- L1 and L2 regularization leave the architecture in place and add different weight-based costs to the loss function.
- Dropout randomly zeros features during training to discourage fragile feature dependence; inference uses scaled outputs without dropping units.
- Choose a strategy by identifying whether you need to change parameter count, weight-based loss pressure, or training-time feature availability.
Key Takeaways
- Overfitting occurs when a model uses its flexibility to memorize training examples instead of learning representations that generalize.
- Model capacity depends on the number of learnable parameters, so reducing model size changes the resources available for memorization.
- L1 and L2 regularization add different weight-based costs to the loss while leaving the architecture in place.
- Dropout randomly removes features during training but uses scaled outputs without dropping units during inference.
- The right intervention depends on whether the problem calls for changing capacity, weight costs, or training-time feature availability.