Neural Network Architecture Design
Hyperparameters are architecture-level choices, such as layer count, units or filters, activation functions, and dropout rates; they are not trained through backpropagation.
Design Choices Before Training
A neural network is shaped by decisions made before training begins. You may choose how many layers it has, how many units or filters belong in each layer, which activation functions to use, whether to use BatchNormalization, and how much dropout to apply. These choices are called hyperparameters. They are architecture-level choices, and they are not trained through backpropagation.
The essential distinction is timing and role: hyperparameters define the model configuration before a training run, while trainable model parameters are the values adjusted as the configured model is trained. Hyperparameter optimization searches over configurations rather than using backpropagation to train those configuration choices.
One Trial at a Time
Hyperparameter optimization is best understood as a sequence of model trials. A search procedure proposes one configuration. A model is constructed with that configuration and trained on the training data. The trained model is then measured on validation data. That validation result becomes feedback for deciding which configuration to test next.
Comparing Two Candidate Designs
Suppose a search procedure is comparing two candidate configurations that differ in architecture choices.
Choose a candidate: The search procedure proposes one set of hyperparameters, such as a selected layer count, unit or filter arrangement, activation functions, and dropout rate.
Build and train: A new model is constructed with those choices and trained on the training data.
Measure validation performance: The trained model is evaluated on validation data. This result is the candidate's feedback for the search.
Compare and continue: The search uses the history of validation results to decide which configuration to evaluate next. The process repeats for another candidate.
A configuration is not judged from its architecture description alone. Its quality is estimated by carrying out a model trial and observing validation performance.
Why Search Costs So Much
A hyperparameter trial is expensive because obtaining its feedback generally requires creating a new model and training it from scratch on the dataset. The search is therefore not merely comparing settings on paper. Every candidate can require another complete model-construction and training effort before its validation performance is known.
This search is typically not guided directly by gradient descent. The source describes the search space as typically non-differentiable and identifies random search, Bayesian optimization, and genetic algorithms as techniques for proposing configurations. Choices such as layer count or activation-function selection are treated as search decisions, while model training is the separate process that produces the validation feedback.
Three Ways to Propose Candidates
The search method determines how the next configuration is proposed from the available choices and the history of earlier validation results. No method removes the cost of evaluating a candidate. The main difference is how each method chooses what to try next.
| Strategy | How it proposes configurations | Important limitation |
|---|---|---|
| Random search | Selects hyperparameter settings at random and evaluates them repeatedly. | It still pays the cost of evaluating every selected candidate. |
| Bayesian optimization | Uses information from earlier evaluations to propose promising settings. | A proposed setting still requires model construction and training. |
| Genetic algorithms | Provide another family of search techniques for proposing configurations. | The search method does not eliminate candidate-evaluation cost. |
The strategies differ in how they propose candidates, not in whether candidates require evaluation.
More sophisticated search is not automatically better in every situation. The source notes that random search can sometimes be the best practical choice despite its simplicity.
When Validation Feedback Becomes a Target
Validation data provides the feedback that drives architecture search, but repeated use of that feedback creates a risk. If hyperparameters are repeatedly updated according to validation performance, the hyperparameters themselves can overfit to the validation data. The validation set is no longer used only to compare independent choices; it becomes part of the evidence shaping the choices.
A Search That Learns the Validation Set
Imagine a search repeatedly choosing the configuration with the strongest validation performance.
First comparisons: Validation data is used to compare the initial candidate configurations.
More rounds: The search keeps using validation results to change which hyperparameters it tries next.
Accumulated influence: The validation set now influences not only the measurement of candidates but also the sequence of architecture choices.
Later assessment: A final test-data measurement remains a distinct later step, rather than another round of repeatedly shaping the search.
A strong validation score after extensive search may partly reflect adaptation to the validation data, which is why a distinct final test measurement is important.
Architecture Choices in Practice
Classify each item as a hyperparameter choice or as part of the model's training process: choosing the number of layers, selecting an activation function, constructing a model with those choices, training the constructed model, and measuring its validation performance.
Hints
- Ask whether the item is selected before the training run or occurs during the trial.
- Remember that validation performance is feedback used to decide what configuration to test next.
Treating layer count, activation functions, or dropout rates as values learned by backpropagation.
These are architecture-level hyperparameters chosen before training, and the source explicitly states that hyperparameters are not trained through backpropagation.
Fix:
Treat them as configuration choices that the search procedure proposes and evaluates through separate model trials.Assuming a candidate architecture can be judged without training it.
The search receives feedback by creating and training a new model, then measuring it on validation data.
Fix:
Include model construction, training, and validation measurement in the evaluation cycle.Assuming a more advanced search method eliminates computational cost.
No search method eliminates the cost of evaluating a candidate.
Fix:
Compare methods by how they propose candidates while remembering that each candidate may require a new training run.Treating repeated validation improvement as completely independent evidence.
Repeated updates based on validation performance can make the hyperparameters overfit to the validation data.
Fix:
Keep a final test-data measurement as a distinct later step in the optimization process.
The Search in One View
- Hyperparameters are architecture-level choices such as layer count, units or filters, activation functions, BatchNormalization use, and dropout rates; they are selected before training and are not trained through backpropagation.
- Hyperparameter optimization repeatedly selects a configuration, constructs a model, trains it, measures validation performance, and uses that result to choose another configuration.
- The process is expensive because each candidate may require a new model and a training run from scratch.
- Random search, Bayesian optimization, and genetic algorithms differ in how they propose candidates, but none removes the cost of candidate evaluation.
- Repeatedly using validation performance to update hyperparameters can overfit the search to the validation set, so a final test-data measurement is a distinct later step.
Key Takeaways
- Hyperparameters define neural network architecture before training, whereas trainable model parameters are adjusted during the model training process.
- Hyperparameter optimization is a repeated cycle of proposing, constructing, training, validating, and selecting.
- Search is computationally expensive because each candidate may require a new model trained from scratch.
- Random search, Bayesian optimization, and genetic algorithms provide different ways to propose configurations in a typically non-differentiable search space.
- Repeated validation-based selection can overfit the validation set, making a distinct final test measurement necessary.