Training Feedforward Neural Networks
A feedforward neural network is a directed acyclic graph whose nodes are neurons and whose directed edges are weighted connections.
The One-Way Route
A feedforward neural network provides a one-way route for information. Values enter through the input layer, pass through neurons connected by weighted links, and eventually reach the output layer. To understand how such a network is trained, first understand the structure that information travels through: neurons arranged in layers and connections that point forward without forming cycles.
Layer Order
The full set of neurons is divided into layer sets named V0 through VT. The input layer is V0. Hidden layers lie between the input and output layers, and the output layer is VT. The layer index records processing order: information starts at V0, moves through any hidden layers, and finishes at VT. These names describe each layer's role in the flow rather than merely giving it a position.
| Layer | Notation | Role in the flow |
|---|---|---|
| Input | V0 | Exposes the starting feature values |
| Hidden | V1 through VT-1 | Performs intermediate transformations |
| Output | VT | Contains the network's final output |
The layer names describe where information is in the processing route.
Input Representation
For an n-dimensional input space, the described input layer contains n feature neurons and one additional constant neuron. Each feature neuron represents one input feature. The constant neuron always outputs 1. Thus, the input layer does more than hold feature values: it also provides this constant value as part of the information available to the next layer.
Neuron Processing
A neuron receives outputs from earlier neurons through weighted connections. It combines those weighted incoming outputs and then processes the result with an activation function. The activation function therefore sits between a neuron's combined incoming information and the output that the neuron passes onward. The exact activation function is not needed to identify its role: it transforms the combined incoming result into the neuron's output.
A Small Network Trace
Following Two Features to One Output
Trace an illustrative network with two feature neurons, one constant neuron, two hidden neurons, and one output neuron.
Start at V0: The input layer contains Feature 1, Feature 2, and the constant neuron. The first two nodes expose the example's feature values, while the constant neuron outputs 1.
Move to the hidden layer: Each hidden neuron receives information through weighted connections from the input layer. The hidden neurons combine their incoming weighted outputs and process the results with activation functions.
Continue forward: The outputs of the hidden neurons travel through weighted connections to the output neuron. The output neuron again combines its incoming weighted outputs and applies its activation function.
Arrive at VT: The output neuron's result is the network's final output at the output layer. No information needs to travel backward or return to an earlier layer in this feedforward route.
The trace is V0 to hidden neurons to VT: feature values and the constant value enter first, hidden neurons perform intermediate transformations, and the output neuron produces the final output.
When tracing a feedforward network, follow the direction of the connections. First identify the values exposed by V0, then track the weighted incoming information received by each hidden neuron, then follow the hidden outputs to VT. The network's graph structure determines this order.
Common Misunderstandings
Treating the network as a loose collection of independent neurons
The neurons and their communication links form a directed graph. The links determine how one stage supplies information to the next.
Fix:
Trace both the neurons and the directed weighted connections that carry information forward.Calling every layer an input or output layer
The input, hidden, and output names describe positions and roles in the complete flow. Hidden layers perform intermediate transformations, while the output layer contains the network's final output.
Fix:
Use V0 for the input layer, V1 through VT-1 for hidden layers, and VT for the output layer.Counting only the feature neurons in the input layer
The described input layer contains n feature neurons plus one constant neuron that always outputs 1.
Fix:
List the feature neurons and the additional constant neuron separately.Treating a neuron's output as just its incoming weighted combination
A neuron processes the combined result with an activation function before producing its output.
Fix:
Include the activation-function step when tracing what a neuron sends onward.Assuming information can move in cycles
A feedforward neural network is a directed acyclic graph, and its information route moves forward through the layers.
Fix:
Follow the directed route from V0 through hidden layers to VT without returning to an earlier stage.
Practice Trace
Describe the information route in a feedforward network whose input layer contains three feature neurons and one constant neuron, followed by one hidden layer and an output layer. Name what happens at the input layer, at each hidden neuron, and at the output neuron.
Hints
- Start with V0 and identify the three feature neurons and the constant neuron.
- Remember that each neuron combines weighted incoming outputs and processes the result with an activation function.
- End at VT, where the network's final output is found.
Key Takeaways
- A feedforward neural network is a directed acyclic graph whose nodes are neurons and whose directed edges are weighted connections.
- The input layer is V0, hidden layers lie between the input and output layers, and the output layer is VT.
- For an n-dimensional input, the described input layer contains n feature neurons plus a constant neuron that always outputs 1.
- A neuron combines weighted incoming outputs and processes the result with an activation function.
- Information moves from the input layer through hidden transformations to the final output layer.
Key Takeaways
- Feedforward neural networks organize neurons and weighted connections as a directed acyclic graph.
- The layer sequence is V0 for input, V1 through VT-1 for hidden layers, and VT for output.
- The input layer represents features and includes a constant neuron that always outputs 1.
- Each neuron combines weighted incoming outputs and applies an activation function.
- Tracing the directed connections reveals how information moves from inputs to the final output.