Concepts / Regularization in Deep Learning

Regularization in Deep Learning

Hyperparameters are architecture-level choices, such as layer count, units or filters, activation functions, and dropout rates; they are not trained through backpropagation.

  • Programming

Choices Before Training

A deep learning model is shaped by decisions made before training begins. You may choose the number of layers, the number of units or filters in each layer, the activation functions, whether to use BatchNormalization, and how much dropout to apply. These choices are hyperparameters. They affect the model, but they are not trained through backpropagation.

shapelearned valuesHyperparameterslayer count, units,activations, dropoutDeep learning modelarchitecture and learnedvaluesTrainable modelparametersupdated during training
What is chosen before training, what is updated by backpropagation, and how do these two kinds of values interact?

A hyperparameter is an architecture-level choice made as part of deciding how a model will be constructed. Examples include layer count, units or filters, activation functions, BatchNormalization usage, and dropout rate. Unlike the values trained inside the model, hyperparameters are not trained through backpropagation.

The Trial-and-Feedback Cycle

Hyperparameter optimization is best understood as a sequence of model trials. A search procedure proposes one configuration. A model is built with that configuration and trained on the training data. The model is then measured on validation data. That validation measurement becomes feedback for deciding which configuration to test next.

proposebuildevaluatefeedbackrepeatSelectconfigurationhyperparameter valuesConstruct modeluse selected architectureTrain modeltraining dataMeasure validationperformancevalidation dataChoose nextconfigurationuse evaluation history
What happens next when a configuration is proposed, trained, evaluated on validation data, and used to choose the next configuration?

Tracing Three Configuration Trials

Suppose a search procedure evaluates several possible model configurations. Trace what the procedure does for each trial.

Propose: The search procedure chooses one configuration containing architecture-level decisions such as layer count, units or filters, activation functions, or dropout rate.

Construct: A new model is created using that configuration.

Train: The model is trained on the training data.

Validate: The trained model is measured on validation data.

Repeat: The validation result becomes feedback for selecting another configuration, and the cycle continues.

A hyperparameter search is a repeated sequence of configuration proposal, model construction, training, validation measurement, and another selection.

Why Trials Cost So Much

A search procedure does not receive useful feedback merely by comparing settings on paper. To discover whether a configuration works well, it generally has to create a new model and train that model from scratch on the dataset. Every candidate configuration can therefore expand into a complete model-training run before its validation performance is known.

constructtrainmeasureproposeConfigurationarchitecture choicesNew modelconstructed from scratchTraining runtraining dataValidation feedbackperformance measurementSearch spacelayer count or activationfunction
How does one hyperparameter trial expand into a full model-training run, and why can ordinary gradients not directly update choices such as layer count or activation function?

Ordinary backpropagation does not directly guide this search because the choices are treated as a typically non-differentiable search space. Layer count and activation-function selection are examples of architecture-level choices, not ordinary trainable values updated by backpropagation. Search methods therefore propose configurations and obtain feedback by evaluating completed trials.

Three Search Strategies

The search method decides what configuration to try next. Random search selects hyperparameter settings at random and evaluates them repeatedly. Bayesian optimization uses information from earlier evaluations to propose promising settings. Genetic algorithms are another family of search techniques. These methods differ in how they use the history of trials, but each still depends on evaluating proposed configurations.

select at randomuse prior resultspropose through a search techniqueRandom searchrandom settingsProposedconfigurationevaluate nextBayesianoptimizationearlier evaluationsGenetic algorithmssearch technique family
How does each search strategy choose its next hyperparameter configuration, and what information does it use from earlier trials?
MethodHow it proposes configurationsUse of earlier evaluations
Random searchSelects hyperparameter settings at randomDoes not use earlier results to make the random selection
Bayesian optimizationProposes promising settingsUses information from earlier evaluations
Genetic algorithmsUses a family of search techniquesThe source identifies it as another search family without specifying a single proposal rule

Do not assume that a more elaborate search method automatically wins. The source notes that random search can sometimes be the best practical choice despite its simplicity, while every method still pays the cost of evaluating candidates.

Validation Feedback and Overfitting

Validation data provides the feedback that drives hyperparameter search. That usefulness creates a risk: if hyperparameters are repeatedly updated according to validation performance, the hyperparameters themselves can overfit to the validation data. The validation set is no longer being used only to compare independent choices; it is also becoming part of the evidence that shapes those choices.

evaluaterepeat selectionshape choicesValidation datacompare choicesHyperparameter trialmodel evaluationValidation datashapes repeated choicesSelectedhyperparameterschosen from validationfeedback
How can repeatedly selecting configurations based on validation performance make the validation set influence the final model too much?

Recognizing Validation-Set Overfitting

A search repeatedly proposes configurations, trains a new model for each one, and keeps using validation performance to decide what to try next. What changes about the role of the validation set?

Initial comparison: The validation data provides feedback for comparing candidate configurations.

Repeated selection: The search repeatedly updates its choices according to validation performance.

Accumulated influence: The validation set becomes part of the evidence shaping the hyperparameters, rather than serving only as a one-time comparison source.

Later measurement: A final test-data measurement remains a distinct later step in the optimization process.

Repeatedly selecting based on validation performance can make the hyperparameters overfit to the validation data, so final evaluation must use a distinct test-data measurement.

Common Search Mistakes

  • Treating hyperparameters as values that backpropagation trains

    These are architecture-level choices and are not trained through backpropagation.

    Fix: Treat them as candidate settings proposed by a hyperparameter search procedure.

  • Counting a configuration comparison as the entire trial

    Validation feedback generally requires creating a new model and training it from scratch.

    Fix: Include model construction, training, and validation measurement in the cost of each candidate.

  • Assuming Bayesian optimization removes evaluation cost

    No search method eliminates the cost of evaluating a candidate.

    Fix: Understand Bayesian optimization as a way to propose promising settings using earlier results, followed by another evaluation.

  • Treating repeated validation selection as independent testing

    Repeated selection can cause the hyperparameters to overfit to the validation data.

    Fix: Use a distinct later test-data measurement for final evaluation.

Practice the Reasoning

MEDIUM

A search procedure evaluates a configuration by constructing a model, training it from scratch, measuring validation performance, and using that result to select the next configuration. Identify which part makes the search expensive, which part prevents ordinary backpropagation from directly choosing the next architecture, and what risk appears after many rounds of using the same validation data.

Hints
  • Separate the cost of obtaining feedback from the method used to propose a configuration.
  • Think about whether architecture-level choices form the same kind of values as those updated through backpropagation.
  • Ask whether repeated validation-based selection leaves the validation set uninvolved.
  1. Hyperparameters shape a model before training and are not trained through backpropagation. Hyperparameter optimization repeatedly proposes configurations, constructs models, trains them, and measures validation performance. Each trial can be expensive because it may require a new model trained from scratch. Random search, Bayesian optimization, and genetic algorithms differ in how they propose configurations, but none removes evaluation cost. Repeatedly choosing hyperparameters from validation results can make the hyperparameters overfit to the validation set, so final test-data measurement must remain distinct.

Key Takeaways

  • Hyperparameters are architecture-level choices such as layer count, units or filters, activation functions, BatchNormalization usage, and dropout rate.
  • Hyperparameters are not trained through backpropagation; search procedures propose them and evaluate the resulting models.
  • Hyperparameter optimization follows a repeated cycle of selection, construction, training, validation measurement, and new selection.
  • Random search, Bayesian optimization, and genetic algorithms are different ways to propose configurations in a typically non-differentiable search space.
  • Repeated validation-based selection can overfit hyperparameters to the validation data, making a distinct final test measurement important.