Learning with Neural Networks
A neural network architecture is specified by (V, E, σ), consisting of the graph and the activation function.
From Structure to Predictor
A neural network can be viewed as a function that maps inputs to outputs. To describe how a neural network learns, separate two decisions. First, specify the network's structure and activation function. Second, assign numerical weights to the network's edges. The first decision fixes the architecture; the second selects a particular predictor within that architecture.
The Architecture Tuple
A neural network architecture is specified by the tuple (V, E, σ). The graph is described by V and E, while σ identifies the activation function. Together, these components specify the fixed structure used by the network before any particular edge weights are chosen.
Describing one architecture
Suppose a generated example uses V to name the network's vertices, E to name its edges, and σ to name its activation function. What does the tuple (V, E, σ) describe?
Identify V: Treat V as the vertex part of the graph.
Identify E: Treat E as the edge part of the graph.
Identify σ: Treat σ as the activation function associated with the architecture.
Combine the parts: The three components together specify the architecture before weights are assigned.
The tuple (V, E, σ) identifies one fixed neural network architecture.
Tracing a Signal Through the Network
The graph provides the network's structure for processing a signal from an input toward an output. The edges connect the graph's vertices, and the activation function is part of the architecture that governs the network's transformation of signals. At this level, the important idea is not a particular numerical calculation but the ordered relationship among the input, the graph, the activation function, and the output.
The architecture determines the available structure for the input-to-output function. The weights have not yet been chosen in this description, so this is an architecture-level view rather than one fully specified hypothesis.
Weights Select One Hypothesis
Weights are assigned to the network edges. Once a particular set of weights w is assigned to the fixed architecture (V, E, σ), the tuple (V, E, σ, w) defines one function, written h_V,E,σ,w. That function is one individual hypothesis: one possible mapping from inputs to outputs represented by the architecture with those particular parameters.
Two weight assignments
Consider one fixed architecture (V, E, σ). In a generated example, assign one collection of edge weights called w_A, then assign a different collection called w_B. What do these assignments represent?
Hold the architecture fixed: The graph and activation function remain the same in both cases.
Choose w_A: The tuple (V, E, σ, w_A) defines one function, h_V,E,σ,w_A.
Choose w_B: The tuple (V, E, σ, w_B) defines another function, h_V,E,σ,w_B.
Compare the results: The two weight assignments select two individual hypotheses while the architecture remains fixed.
The weight assignment is a parameter choice that identifies an individual hypothesis; it is not an additional structural choice in the fixed architecture.
Building a Hypothesis Class
A hypothesis class is formed by keeping the architecture fixed and allowing the weights to vary. The notation H_V,E,σ denotes the collection of functions obtained from all possible choices of weights for that architecture. Each allowed weight choice gives one predictor, and the collection of those predictors is the neural-network hypothesis class.
| Choice | Fixed or variable? | Role |
|---|---|---|
| Graph | Fixed | Part of the architecture |
| Activation function σ | Fixed | Part of the architecture |
| Edge weights w | Variable | Parameters of an individual hypothesis |
| Functions in H_V,E,σ | Collected | All predictors produced by varying the weights |
Reading the Hypothesis-Class Notation
Read H_V,E,σ as a class indexed by one fixed architecture. The subscripts V, E, and σ identify the graph and activation function that remain fixed. The weights do not appear as fixed subscripts because they are allowed to vary across the functions in the class. For one selected weight assignment, the more specific notation is h_V,E,σ,w.
Common Modeling Mistakes
Treating the weights as part of the fixed architecture.
The architecture is specified by (V, E, σ), while weights are parameters used to select an individual hypothesis.
Fix:
Describe the architecture first, then add a weight assignment to obtain (V, E, σ, w).Confusing one hypothesis with the hypothesis class.
One weight assignment defines h_V,E,σ,w; H_V,E,σ contains the functions obtained by varying the weights.
Fix:
Use h_V,E,σ,w for one selected function and H_V,E,σ for the collection.Assuming a new weight assignment changes the graph.
Changing weights can select a different function while the graph and activation function remain fixed.
Fix:
Keep (V, E, σ) fixed and describe w_A and w_B as different parameter choices.Leaving the activation function out of the architecture description.
The architecture is specified by (V, E, σ), which includes the activation function.
Fix:
Include σ when naming the complete architecture.
Practice the Separation
A neural network keeps V, E, and σ unchanged but uses two different edge-weight assignments, w_1 and w_2. Identify the notation for each individual hypothesis and the notation for the hypothesis class containing both.
Hints
- An individual hypothesis includes a particular weight assignment.
- The hypothesis class keeps the architecture in its subscript and allows the weights to vary.
Checking the notation
Use the fixed architecture (V, E, σ) and the two weight assignments w_1 and w_2.
Name the first hypothesis: Attach w_1 to the fixed architecture: h_V,E,σ,w_1.
Name the second hypothesis: Attach w_2 to the fixed architecture: h_V,E,σ,w_2.
Name the class: Use H_V,E,σ for the collection of functions produced by varying the weights.
The two individual hypotheses are h_V,E,σ,w_1 and h_V,E,σ,w_2, and both belong to H_V,E,σ.
Key Takeaways
- A neural network architecture is specified by (V, E, σ): the graph components and the activation function.
- Weights are assigned to network edges and act as parameters of an individual hypothesis.
- The tuple (V, E, σ, w) defines one function, h_V,E,σ,w.
- Fixing the architecture and varying the weights produces the hypothesis class H_V,E,σ.
- The architecture is structural; the weights are variable parameters that select functions within the class.
Key Takeaways
- The tuple (V, E, σ) specifies a neural network architecture through its graph and activation function.
- A particular edge-weight assignment w completes the architecture into one hypothesis h_V,E,σ,w.
- The hypothesis class H_V,E,σ contains the functions obtained by varying weights while keeping the architecture fixed.
- Weights change the selected function, not the structural definition of the fixed architecture.