Concepts / Feedforward artificial neural networks

Feedforward artificial neural networks

A sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

  • Programming
Interactive lab

Try it: Neural Network Forward Pass

How a neural network turns inputs into an output: each neuron computes a weighted sum plus bias, applies an activation function, and passes the result on.

How it works

  1. Each hidden neuron computes z = w1·x1 + w2·x2 + b.
  2. It applies an activation: ReLU keeps positive values and zeroes negatives; sigmoid squashes z into 0–1.
  3. The output neuron does the same with the hidden activations as its inputs.
  4. Changing any input, weight or bias changes every value downstream of it.

Default run (7 steps): Inputs x1 = 1, x2 = 0.5. Activation: ReLU. … Prediction y = ReLU(1.25) = max(0, 1.25) = 1.25.

Simplified: Tiny 2-2-1 network with hand-set weights, not a trained model. No training happens here. Activations offered: ReLU, sigmoid and sign (texts differ on the value of sign at exactly 0; here it is 0).

Educational simulation

Loading the simulation…

From Inputs to Outputs

A feedforward artificial neural network transforms an input into an output through connected layers. Its structure includes an input layer, one or more hidden layers, and an output layer. Connections between units carry real-valued weights. The defining structural feature is that the connections contain no loops: information moves through the network without a unit's output being able to influence its own input.

weighted connectionsweighted connectionsdescribesdescribesInput layerreceives inputHidden layerinternal unitsOutput layerproduces outputReal-valued weighton a connectionReal-valued weighton a connection
How does information move through the layers, and where do the real-valued weights belong?

A Single Hidden Layer as a Function Approximator

The universal approximation property states that a sufficiently large feedforward network with one hidden layer of sigmoid units can approximate any continuous function on a compact input region to any desired accuracy. This is a capacity statement about what an appropriate network can approximate. It is not a claim that every network can approximate every continuous function, nor that one fixed network automatically has this ability.

can approximatetoproperty does not sayContinuous functioncompact input regionSufficiently largenetworkone sigmoid hidden layerDesired accuracyapproximationEvery individualnetworknot claimedOne fixed networknot claimed
What does the universal approximation property claim, and what does it not claim?

Interpreting the capacity claim

Suppose a continuous target function is defined over a compact input region. What would the universal approximation property allow you to say?

Identify the target: The target is a continuous function over a compact input region.

Choose the network form: Consider a feedforward network with one hidden layer whose units use sigmoid activation functions.

Apply the property: The property says that a sufficiently large network of this form can approximate the target function to any desired accuracy.

Avoid overclaiming: The statement does not say that every network, every hidden-layer size, or one fixed network can approximate every such function.

The universal approximation property is an existence and capacity claim: an appropriately sufficiently large single-hidden-layer network can provide the requested approximation.

Why Nonlinearity Matters

The hidden layer does not expand a network's representational ability merely because it adds another layer. The crucial ingredient is nonlinearity in the activation functions. An activation function determines how a unit responds to the activation pattern arriving from the network's input units. The universal approximation result depends on nonlinear activation functions; sigmoid units are one stated example, and other nonlinear activation functions can work when they satisfy mild conditions.

A network with no hidden layers can represent only a very small fraction of possible input-output functions. Adding a hidden layer with suitable nonlinear units greatly expands the class of functions the network can approximate. The important contrast is therefore not simply shallow versus deep. It is linear transformations alone versus transformations that include nonlinear activation behavior.

producesenablesLinear activationsmall function classLinear transformationeven across layersNonlinearactivationsigmoid is one exampleBroader approximationabilitywith suitable hidden units
What changes when nonlinear activation functions are added instead of using only linear transformations?

The Linear-Layer Limitation

If every activation function in a multi-layer network is linear, composing one layer with the next still produces a linear function. Consequently, several layers with only linear activations do not create a richer class of functions than a network with no hidden layers. The entire multi-layer network is equivalent in representational capability to a network with no hidden layers.

linear mappinglinear mappingcomposes toequivalent capability toInputnetwork inputLinear layer 1linear activationLinear layer 2linear activationOne lineartransformationequivalent capabilityNo-hidden-layernetworksame limitation
How do multiple layers with only linear activations combine, and why do they retain the capability of a network with no hidden layers?

Evaluating a proposed deep network

A proposed network has several layers, but every unit uses a linear activation function. Does the layer count alone give it the broader approximation ability associated with the universal approximation property?

Inspect the activations: Every activation function is linear.

Trace the composition: The output of one linear layer becomes the input to another linear layer, so the layers compose into a linear function.

Compare with the baseline: The resulting network has the same representational capability as a network with no hidden layers.

State the conclusion: The network does not gain the relevant broader approximation ability from depth alone.

Several linear layers remain equivalent to one linear transformation, so nonlinearity is essential for the expanded approximation capability.

Recognizing Feedforward Structure

The input layer receives the input, the output layer produces the output, and any layer between them is a hidden layer. A feedforward ANN may have one or more hidden layers. Hidden means internal to the network; it does not mean disconnected or inactive. Units are linked by connections, and every connection has an associated real-valued weight.

To classify the network, inspect the connection pattern rather than merely observing that information enters at one side and eventually produces an output. A feedforward network has no loops. If a unit's output can influence its own input through a path in the network, the network contains a loop and is recurrent instead.

forward connectionforward connectionconnectionconnectionloopInput layerHidden layerOutput layerno loopInput layerHidden layerOutput layerloop present
How can you classify a network by checking whether its connections contain a loop?

Classification Practice

EASY

Consider a network organized as input layer, hidden layer, and output layer. Every connection has a real-valued weight. The connections move from the input layer to the hidden layer and from the hidden layer to the output layer, with no path returning to an earlier unit. Classify the network and explain which structural feature determines your answer.

Hints
  • Check whether any connection pattern forms a loop.
  • The weights are part of the network specification, but they do not determine whether the network is feedforward.
  • Use the distinction between feedforward and recurrent networks.
MEDIUM

A second network has the same input, hidden, and output layers, but one connection creates a loop through which a unit's output can influence its own input. Classify this network and explain why the layer names alone are not enough.

Hints
  • The presence of a loop is decisive.
  • A network can still have input, hidden, and output layers while being recurrent.

Key Takeaways

  1. A sufficiently large single-hidden-layer feedforward network with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.
  2. This is a capacity claim about an appropriately sufficiently large network, not a claim about every network or one fixed network.
  3. Nonlinear activation functions provide the crucial approximation ability beyond the small function class available without hidden layers.
  4. Several layers with only linear activation functions remain equivalent to one linear transformation and do not overcome the limitation.
  5. A feedforward ANN has input, hidden, and output layers, real-valued weights on its connections, and no loops. A single loop makes the network recurrent.

Key Takeaways

  • The universal approximation property concerns what a sufficiently large single-hidden-layer network with suitable nonlinear units can approximate.
  • Nonlinearity, not layer count alone, expands the class of functions a network can represent.
  • Stacking linear layers still produces a linear function, so it retains the capability of a network with no hidden layers.
  • Feedforward networks are defined by the absence of loops in their connections.
  • Input, hidden, and output layers describe a network's organization, while real-valued weights describe its connections.