Concepts / Neural Network Architecture Design

Neural Network Architecture Design

Hyperparameters are architecture-level choices, such as layer count, units or filters, activation functions, and dropout rates; they are not trained through backpropagation.

  • Programming

Design Choices Before Training

A neural network is shaped by decisions made before training begins. You may choose how many layers it has, how many units or filters belong in each layer, which activation functions to use, whether to use BatchNormalization, and how much dropout to apply. These choices are called hyperparameters. They are architecture-level choices, and they are not trained through backpropagation.

includesincludesincludesincludesupdated throughHyperparameterschosen before trainingLayer countarchitecture choiceModel trainingbackpropagation does nottrain hyperparametersUnits or filtersarchitecture choiceTrainable modelparametersadjusted during trainingActivation functionsarchitecture choiceDropout ratesarchitecture choice
Which values are selected before training, and which values are learned during the model's training run?

The essential distinction is timing and role: hyperparameters define the model configuration before a training run, while trainable model parameters are the values adjusted as the configured model is trained. Hyperparameter optimization searches over configurations rather than using backpropagation to train those configuration choices.

One Trial at a Time

Hyperparameter optimization is best understood as a sequence of model trials. A search procedure proposes one configuration. A model is constructed with that configuration and trained on the training data. The trained model is then measured on validation data. That validation result becomes feedback for deciding which configuration to test next.

build from choicerun trainingevaluateprovide feedbackrepeat with another candidateSelect configurationcandidate hyperparametersConstruct modeluse the configurationTrain modeltraining dataMeasure validationperformancevalidation dataChoose nextconfigurationuse validation feedback
What happens after a candidate architecture is selected, and how does its validation result affect the next candidate?

Comparing Two Candidate Designs

Suppose a search procedure is comparing two candidate configurations that differ in architecture choices.

Choose a candidate: The search procedure proposes one set of hyperparameters, such as a selected layer count, unit or filter arrangement, activation functions, and dropout rate.

Build and train: A new model is constructed with those choices and trained on the training data.

Measure validation performance: The trained model is evaluated on validation data. This result is the candidate's feedback for the search.

Compare and continue: The search uses the history of validation results to decide which configuration to evaluate next. The process repeats for another candidate.

A configuration is not judged from its architecture description alone. Its quality is estimated by carrying out a model trial and observing validation performance.

Why Search Costs So Much

A hyperparameter trial is expensive because obtaining its feedback generally requires creating a new model and training it from scratch on the dataset. The search is therefore not merely comparing settings on paper. Every candidate can require another complete model-construction and training effort before its validation performance is known.

specifiesrequiresproducesCandidateconfigurationhyperparameter choicesNew modelconstructed from scratchTraining runtraining dataValidation feedbackperformance measurement
How does one proposed configuration expand into a full training run before the search receives validation feedback?

This search is typically not guided directly by gradient descent. The source describes the search space as typically non-differentiable and identifies random search, Bayesian optimization, and genetic algorithms as techniques for proposing configurations. Choices such as layer count or activation-function selection are treated as search decisions, while model training is the separate process that produces the validation feedback.

Three Ways to Propose Candidates

The search method determines how the next configuration is proposed from the available choices and the history of earlier validation results. No method removes the cost of evaluating a candidate. The main difference is how each method chooses what to try next.

selects at randomuses earlier results to propose promising settingssearches as a family of techniquesRandom searchrandom settingsNext configurationmust still be evaluatedBayesianoptimizationearlier evaluationsGenetic algorithmscandidate populations
How do random search, Bayesian optimization, and genetic algorithms choose their next candidate designs?
StrategyHow it proposes configurationsImportant limitation
Random searchSelects hyperparameter settings at random and evaluates them repeatedly.It still pays the cost of evaluating every selected candidate.
Bayesian optimizationUses information from earlier evaluations to propose promising settings.A proposed setting still requires model construction and training.
Genetic algorithmsProvide another family of search techniques for proposing configurations.The search method does not eliminate candidate-evaluation cost.

The strategies differ in how they propose candidates, not in whether candidates require evaluation.

More sophisticated search is not automatically better in every situation. The source notes that random search can sometimes be the best practical choice despite its simplicity.

When Validation Feedback Becomes a Target

Validation data provides the feedback that drives architecture search, but repeated use of that feedback creates a risk. If hyperparameters are repeatedly updated according to validation performance, the hyperparameters themselves can overfit to the validation data. The validation set is no longer used only to compare independent choices; it becomes part of the evidence shaping the choices.

reuse validation feedbackcan gradually shapeevaluate later on distinct dataInitial searchvalidation comparescandidatesRepeated selectionchoices shaped byvalidation resultsSearch processspecialized to validationdataFinal testmeasurementdistinct later step
How can repeated selection using validation scores make the search process specialize to the validation set?

A Search That Learns the Validation Set

Imagine a search repeatedly choosing the configuration with the strongest validation performance.

First comparisons: Validation data is used to compare the initial candidate configurations.

More rounds: The search keeps using validation results to change which hyperparameters it tries next.

Accumulated influence: The validation set now influences not only the measurement of candidates but also the sequence of architecture choices.

Later assessment: A final test-data measurement remains a distinct later step, rather than another round of repeatedly shaping the search.

A strong validation score after extensive search may partly reflect adaptation to the validation data, which is why a distinct final test measurement is important.

Architecture Choices in Practice

shapesshapesaffectscan be included inaffectsNeural networkresulting modelLayer countstructureUnits or filterslayer sizeActivation functionslayer behaviorBatchNormalizationoptional design choiceDropout rateregularization choice
How do common architecture choices connect to the structure and behavior of a neural network?
EASY

Classify each item as a hyperparameter choice or as part of the model's training process: choosing the number of layers, selecting an activation function, constructing a model with those choices, training the constructed model, and measuring its validation performance.

Hints
  • Ask whether the item is selected before the training run or occurs during the trial.
  • Remember that validation performance is feedback used to decide what configuration to test next.
  • Treating layer count, activation functions, or dropout rates as values learned by backpropagation.

    These are architecture-level hyperparameters chosen before training, and the source explicitly states that hyperparameters are not trained through backpropagation.

    Fix: Treat them as configuration choices that the search procedure proposes and evaluates through separate model trials.

  • Assuming a candidate architecture can be judged without training it.

    The search receives feedback by creating and training a new model, then measuring it on validation data.

    Fix: Include model construction, training, and validation measurement in the evaluation cycle.

  • Assuming a more advanced search method eliminates computational cost.

    No search method eliminates the cost of evaluating a candidate.

    Fix: Compare methods by how they propose candidates while remembering that each candidate may require a new training run.

  • Treating repeated validation improvement as completely independent evidence.

    Repeated updates based on validation performance can make the hyperparameters overfit to the validation data.

    Fix: Keep a final test-data measurement as a distinct later step in the optimization process.

The Search in One View

  1. Hyperparameters are architecture-level choices such as layer count, units or filters, activation functions, BatchNormalization use, and dropout rates; they are selected before training and are not trained through backpropagation.
  2. Hyperparameter optimization repeatedly selects a configuration, constructs a model, trains it, measures validation performance, and uses that result to choose another configuration.
  3. The process is expensive because each candidate may require a new model and a training run from scratch.
  4. Random search, Bayesian optimization, and genetic algorithms differ in how they propose candidates, but none removes the cost of candidate evaluation.
  5. Repeatedly using validation performance to update hyperparameters can overfit the search to the validation set, so a final test-data measurement is a distinct later step.

Key Takeaways

  • Hyperparameters define neural network architecture before training, whereas trainable model parameters are adjusted during the model training process.
  • Hyperparameter optimization is a repeated cycle of proposing, constructing, training, validating, and selecting.
  • Search is computationally expensive because each candidate may require a new model trained from scratch.
  • Random search, Bayesian optimization, and genetic algorithms provide different ways to propose configurations in a typically non-differentiable search space.
  • Repeated validation-based selection can overfit the validation set, making a distinct final test measurement necessary.