Concepts / Introduction to Feedforward Neural Networks

Introduction to Feedforward Neural Networks

A neural network architecture is specified by (V, E, σ), consisting of the graph and the activation function.

  • Programming
Interactive lab

Try it: Neural Network Forward Pass

How a neural network turns inputs into an output: each neuron computes a weighted sum plus bias, applies an activation function, and passes the result on.

How it works

  1. Each hidden neuron computes z = w1·x1 + w2·x2 + b.
  2. It applies an activation: ReLU keeps positive values and zeroes negatives; sigmoid squashes z into 0–1.
  3. The output neuron does the same with the hidden activations as its inputs.
  4. Changing any input, weight or bias changes every value downstream of it.

Default run (7 steps): Inputs x1 = 1, x2 = 0.5. Activation: ReLU. … Prediction y = ReLU(1.25) = max(0, 1.25) = 1.25.

Simplified: Tiny 2-2-1 network with hand-set weights, not a trained model. No training happens here. Activations offered: ReLU, sigmoid and sign (texts differ on the value of sign at exactly 0; here it is 0).

Educational simulation

Loading the simulation…

From Structure to Prediction

A neural network can be viewed as a function that maps inputs to outputs. To describe a collection of possible neural-network predictors, separate two decisions: first choose the network structure, and then choose the numerical weights assigned to its edges. The structure determines the architecture; the weights select one particular function within that architecture.

The central distinction is simple: the architecture is the fixed structure, while the weights are the adjustable parameters that select an individual hypothesis.

Architecture as a Three-Part Description

A neural network architecture is specified by the tuple (V, E, σ). The graph contributes its vertices V and directed edges E. The activation function contributes σ. Together, these describe the structure and activation behavior that are held fixed when defining the architecture.

containscontainsusesArchitectureVverticesEdirected edgesσactivation function
What does each part of (V, E, σ) contribute to a neural-network architecture?

Reading an Architecture Description

Suppose an architecture is described as (V, E, σ). Identify which part represents the graph and which part represents the activation function.

Identify the graph: The graph is represented by V and E: V denotes the vertices and E denotes the directed edges.

Identify the activation function: The symbol σ denotes the activation function.

Keep the categories separate: V, E, and σ together specify the architecture. Numerical edge weights are not part of this three-part architecture description.

The architecture consists of the graph described by V and E together with the activation function σ.

Tracing a Network Function

To understand the network as a function, follow the directed structure from an input toward an output. The graph specifies which vertices are connected by directed edges, and the activation function is part of the computation associated with the architecture. The result is a mapping from inputs to outputs.

entersfollowsusesproducesInputVertexDirected edgeσactivation functionOutput
How does an input move through the network description toward an output?

Weights Select One Hypothesis

Weights are assigned to the network edges. They act as the parameters of an individual hypothesis. Once the architecture (V, E, σ) and a particular assignment of weights w are fixed, the tuple (V, E, σ, w) defines one function, written h_V,E,σ,w.

is processed byparameterizeswith w definesInputwedge weights(V, E, σ)fixed architectureh_V,E,σ,wone function
How do numerical weights assigned to edges determine one particular hypothesis?

Comparing Two Weight Assignments

Consider one fixed architecture (V, E, σ). Compare two possible weight assignments, w₁ and w₂.

Hold the architecture fixed: The vertices V, directed edges E, and activation function σ remain the same in both cases.

Change the weights: The first network uses w₁, while the second uses w₂.

Name the resulting functions: The first choice defines h_V,E,σ,w₁, and the second defines h_V,E,σ,w₂.

The two weight assignments can specify two different hypotheses even though the architecture is unchanged.

Building the Hypothesis Class

A hypothesis class is formed by fixing the architecture and allowing the edge weights to vary. Every possible choice of weights gives one predictor, and the collection of all such predictors is the hypothesis class H_V,E,σ. The subscript records the architecture components that stay fixed; the varying weights are represented by the individual functions contained in the class.

combine withcombine withdefinesdefinesbelongs tobelongs to(V, E, σ)w₁h_V,E,σ,w₁H_V,E,σall resulting functionsw₂h_V,E,σ,w₂
What changes when the architecture stays fixed but the edge weights vary?
ItemWhat happens to it?Role
VHeld fixedVertices in the graph
EHeld fixedDirected edges in the graph
σHeld fixedActivation function
wAllowed to varyWeights assigned to network edges
h_V,E,σ,wChanges when w changesOne individual hypothesis
H_V,E,σCollects the resulting functionsHypothesis class

Reading the Class Notation

Read H_V,E,σ as “the hypothesis class for the architecture specified by V, E, and σ.” Read h_V,E,σ,w as “the individual function obtained from that architecture with the particular weight assignment w.” The notation makes the modeling decision visible: V, E, and σ identify the fixed architecture, while w identifies one choice within the class.

indexed byindexed byindexed bynot part of class subscriptHhypothesis classVverticesEdirected edgesσactivation functionwvarying edge weights
How do the architecture, activation function, and weight assignments map onto the notation?

The absence of w from the subscript of H_V,E,σ is meaningful: the class keeps the architecture fixed while collecting the functions produced by varying w.

Mistakes Beginners Make

  • Treating the weights as part of the fixed architecture

    The architecture is specified by (V, E, σ). Adding w identifies one individual hypothesis based on that architecture.

    Fix: Keep (V, E, σ) fixed when defining the architecture, then use w to select h_V,E,σ,w.

  • Confusing one hypothesis with the hypothesis class

    h_V,E,σ,w denotes one function for one particular weight assignment. H_V,E,σ denotes the collection obtained by varying the weights.

    Fix: Use h for an individual predictor and H for the collection of predictors.

  • Assuming a changed weight creates a changed graph

    The important modeling distinction is that weights can select a different function while the graph and activation function remain fixed.

    Fix: Ask whether the vertices, edges, or activation function changed. If only w changed, the architecture is still the same.

  • Ignoring the activation function in the architecture description

    The architecture is specified by the graph together with the activation function σ.

    Fix: Include all three components: V, E, and σ.

When reading a neural-network description, classify every symbol as either structural or parameter-related. V, E, and σ describe the fixed architecture; w describes the edge-weight assignment used to obtain one hypothesis.

Check Your Understanding

EASY

A neural-network model uses a fixed architecture (V, E, σ). You are given two weight assignments, w₁ and w₂. State what remains fixed, name the two individual hypotheses, and identify the hypothesis class that contains them.

Hints
  • The architecture is the part described by V, E, and σ.
  • Attach each weight assignment to the architecture to name one hypothesis.
  • The hypothesis class keeps the architecture in its subscript and collects the resulting functions.

What do you think happens?

If the graph and activation function stay fixed but the edge weights change, do you still have the same individual hypothesis?

  • Yes, because the architecture is unchanged
  • No, because the weight assignment selects the individual hypothesis
  • No, because the graph must also change
  • Yes, because weights are not part of a neural network
Reveal answer

Answer: No, because the weight assignment selects the individual hypothesis.

The fixed architecture can support multiple hypotheses. Each possible weight assignment w defines an individual function h_V,E,σ,w, and the collection of those functions forms H_V,E,σ.

Summary

  1. A neural-network architecture is specified by (V, E, σ): vertices, directed edges, and an activation function.
  2. Weights are assigned to network edges and act as parameters of an individual hypothesis.
  3. The tuple (V, E, σ, w) defines one function h_V,E,σ,w.
  4. Fixing V, E, and σ while varying w produces the hypothesis class H_V,E,σ.
  5. The architecture determines the fixed structure; the weight assignment selects one function within that structure.

Key Takeaways

  • The architecture of a feedforward neural network is represented by (V, E, σ).
  • The graph and activation function describe the fixed structure, while edge weights are parameters.
  • One weight assignment produces one hypothesis h_V,E,σ,w.
  • Varying the weights while keeping the architecture fixed produces the hypothesis class H_V,E,σ.
  • The notation distinguishes structural choices from the parameters that select an individual predictor.