Model Training and Backpropagation
Hyperparameters are architecture-level choices, such as layer count, units or filters, activation functions, and dropout rates; they are not trained through backpropagation.
The Choices Before Training
Training a deep learning model involves two different kinds of decisions. Some choices describe the model before training begins, such as the number of layers, the number of units or filters, the activation functions, whether to use BatchNormalization, and the amount of dropout. These are hyperparameters. Other model values are learned during model training. Hyperparameters are not trained through backpropagation; instead, they are selected by a search process that repeatedly tests complete model configurations.
One Hyperparameter Trial
Hyperparameter optimization is best understood as a sequence of model trials. A search procedure proposes one configuration. A model is built using that configuration and trained on the training data. The trained model is then measured on validation data. That validation result becomes feedback for selecting the next configuration. The search repeats this cycle until it identifies a configuration that performs well according to the validation measurements.
Comparing Two Candidate Configurations
A search procedure must compare two possible model configurations. What does it have to do before it can use validation performance to choose between them?
Propose the first configuration: The search procedure selects one set of hyperparameter choices.
Construct and train: A new model is created with that configuration and trained on the training data.
Measure validation performance: The trained model is evaluated on validation data, producing feedback about the configuration.
Repeat for the second configuration: The same construction, training, and validation process is carried out for the other candidate.
Compare the feedback: The validation measurements help determine which configuration should be considered more promising.
The search compares configurations through completed model trials, not merely by comparing their written settings.
Backpropagation’s Place
Backpropagation belongs to model training, whereas hyperparameter optimization chooses among model configurations. A configuration can determine the layer count, units or filters, activation functions, and dropout rate before training starts. Backpropagation does not directly train those architecture-level choices. Instead, the search procedure changes the configuration between trials, constructs a model for each choice, and obtains feedback through training and validation.
Why Search Costs So Much
Validation feedback is expensive because discovering whether a configuration works well generally requires creating a new model and training it from scratch on the dataset. A search therefore pays for model construction and training before it receives the validation result. If many configurations are tested, this expense is repeated many times.
The search space is typically non-differentiable for choices such as layer count or activation function. These choices are alternatives in model construction rather than ordinary continuously adjusted training values. Consequently, hyperparameter search is generally handled by proposing and evaluating configurations instead of obtaining direct gradients for those choices through backpropagation.
Three Search Strategies
| Strategy | How the next configuration is proposed | Use of previous results |
|---|---|---|
| Random search | Selects hyperparameter settings at random and evaluates them repeatedly | Does not depend on a sophisticated model of earlier results |
| Bayesian optimization | Proposes promising settings | Uses information from earlier evaluations |
| Genetic algorithms | Uses a genetic-algorithm search approach to propose candidates | Belongs to the family of search techniques used for configuration spaces |
When Validation Becomes a Target
Validation data provides the feedback that drives hyperparameter search, but repeated use of that feedback creates a risk. If hyperparameters are repeatedly updated according to validation performance, the hyperparameters themselves can overfit to the validation data. The validation set then influences the choices too closely instead of serving only as an independent comparison for each choice.
Treat the final test-data measurement as a distinct later step. Validation performance guides the search, while the test measurement is used afterward rather than repeatedly steering hyperparameter choices.
Mistakes to Avoid
Treating layer count, activation functions, or dropout rates as values that backpropagation will learn automatically.
These are architecture-level hyperparameters chosen before training and are not trained through backpropagation.
Fix:
Use a hyperparameter search procedure to propose configurations, then train and validate each proposed model.Assuming that comparing hyperparameter settings on paper is enough.
Validation feedback generally requires a new model and a training run.
Fix:
Count model construction and training as part of every candidate evaluation.Assuming that Bayesian optimization or genetic algorithms make evaluation free.
No search method eliminates the cost of evaluating a candidate.
Fix:
Separate the method for proposing configurations from the required training and validation work.Treating the best validation result as an untouched final measurement after many search decisions.
Repeated selection can make hyperparameters overfit to the validation data.
Fix:
Use a distinct later test-data measurement for final evaluation.
Check Your Understanding
A search procedure proposes a configuration with a different activation function. Explain why ordinary backpropagation does not directly choose that activation function, list the steps needed to obtain feedback about the configuration, and identify which later measurement should remain distinct from the validation-guided search.
Hints
- Start by identifying whether the activation function is an architecture-level choice or a value learned during training.
- Recall the sequence: configuration selection, model construction, training, validation measurement, and new selection.
- The final measurement is not another feedback loop for repeatedly changing hyperparameters.
What do you think happens?
A search method has already evaluated several configurations. Does selecting a more sophisticated proposal method remove the need to train the next proposed model?
Reveal answer
Answer: No, because the candidate still needs model construction and training before validation feedback is available.
Search methods differ in how they propose configurations, but evaluating a candidate still generally requires creating and training a new model before validation performance can be measured.
Key Takeaways
- Hyperparameters are architecture-level choices made before training, including layer count, units or filters, activation functions, and dropout rates; they are not trained through backpropagation.
- Hyperparameter optimization repeats selection, model construction, training, validation measurement, and new selection.
- Each candidate can be expensive because it generally requires a new model and training from scratch.
- Random search, Bayesian optimization, and genetic algorithms differ mainly in how they propose the next configuration, not in whether candidates must be evaluated.
- Repeatedly using validation performance to change hyperparameters can overfit the validation set, so a distinct later test-data measurement is important.
Key Takeaways
- Hyperparameters define important architecture choices before training and are not learned through backpropagation.
- Hyperparameter optimization is a repeated cycle of proposing, constructing, training, validating, and selecting.
- Search is computationally expensive because each candidate generally requires a new training run.
- Random search, Bayesian optimization, and genetic algorithms provide different ways to propose configurations in a typically non-differentiable search space.
- Repeated validation-guided selection can overfit the validation set, making a distinct final test measurement necessary.