Concepts / Function approximation with neural networks

Function approximation with neural networks

A sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

  • Programming

From simple mappings to approximation

A neural network can be viewed as a system that maps inputs to outputs. The important question is not merely whether a network produces an output, but what kinds of input-output functions it can represent or approximate. A network with no hidden layers can represent only a very small fraction of possible input-output functions. The universal approximation property is significant because it describes how a sufficiently wide hidden layer with suitable nonlinear units greatly expands that capability.

The central claim is about approximation capacity: a sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

Tracing a target function

Suppose a continuous target function is defined over a bounded input region. The universal approximation property says that a feedforward network with one hidden layer can approximate that target as closely as desired when the hidden layer is sufficiently large and its units use sigmoid activation functions. The result does not say that one small network automatically matches the target. It says that a network with enough capacity exists within this architectural family.

approximated byproducesContinuous targetbounded input regionSingle hidden layersufficiently wideApproximationdesired accuracy
How can a sufficiently wide single-hidden-layer network approximate a continuous target function over a bounded input region?

Reading the approximation claim

Consider a continuous target function defined on a bounded input region. What does the universal approximation property allow us to claim about a feedforward network with one hidden layer?

Identify the target: The target is a continuous function whose inputs are restricted to a compact, or bounded, region.

Choose the stated architecture: Use a feedforward network with one hidden layer containing sigmoid units.

Apply the capacity condition: The hidden layer must be sufficiently large. The property is not presented as a guarantee for an arbitrarily small network.

State the result: With sufficient size, the network can approximate the target function to any desired accuracy.

The property establishes the existence of a sufficiently large single-hidden-layer network that can approximate the continuous target on the compact input region to any desired accuracy.

Why activation nonlinearity matters

An activation function determines how a unit responds to the activation pattern arriving from the network's input units. The universal approximation result depends on using nonlinear activation functions. Sigmoid units are one stated example, and other nonlinear activation functions can work when they satisfy mild conditions. Nonlinearity is therefore the ingredient that allows hidden units to contribute more than a purely linear input-output relationship.

Adding a hidden layer by itself is not enough. A hidden layer helps expand the class of functions only when its units introduce nonlinearity. This is why the phrase single hidden layer with sigmoid units is important: both the hidden layer and the nonlinear activation behavior are part of the approximation claim.

linear transformationtransformedcontributes toInputlinear pathOutput mappinglinearInputnonlinear pathNonlinear activationsigmoid or other suitablefunctionOutput mappingricher function class
What changes in the input-output mapping when nonlinear activation functions are inserted between layers instead of using only linear transformations?

Tracing a linear stack

Now remove the nonlinear behavior and let every activation function be linear. The network may still contain several layers, but each layer performs a linear transformation. When linear functions are composed with other linear functions, the result remains linear. Consequently, the entire multi-layer network is equivalent to a network with no hidden layers for this purpose.

passes throughpasses throughcomposes toInputinput valuesLinear layer 1linearLinear layer 2linearLinear mappingequivalent to one linearlayer
What function does a stack of linear layers compute after the layers are composed, and why is it no more expressive than one linear layer?

Several layers, unchanged function class

A feedforward network has several layers, but every unit uses a linear activation function. Does the number of layers by itself make the network capable of representing a richer class of functions?

Inspect each layer: Every layer applies a linear transformation because every activation function is linear.

Compose the layers: The output of one linear transformation becomes the input to the next, and linear functions composed with linear functions remain linear.

Compare with the baseline: The complete stack is equivalent to a network with no hidden layers for the resulting function class.

Several linear layers do not overcome the limitation of a network with no hidden layers. Nonlinearity, rather than layer count alone, is the crucial ingredient.

Capacity is not a guarantee for every network

The wording sufficiently large is essential. The universal approximation property is a capacity claim: it says that a suitable network of sufficient size can approximate any continuous function on the stated compact input region to any desired accuracy. It does not say that every single-hidden-layer network can do so, that every width is sufficient, or that an arbitrary network will automatically approximate a chosen target.

can approximatedoes not guaranteeUniversalapproximationpropertysome sufficiently largenetworkContinuous targetcompact input regionIndividual networksize not specifiedAutomaticapproximationnot implied
What is the difference between saying that some sufficiently large network can approximate a target and saying that every network can do so?
  • Interpreting universal approximation to mean that every network can approximate every continuous function.

    The result applies to a sufficiently large network, not to every individual network.

    Fix: State the claim as the existence of a sufficiently large suitable network with one hidden layer.

  • Assuming that adding hidden layers is enough to create complex function representations.

    Composing linear functions with linear functions still produces a linear function.

    Fix: Check whether the units introduce nonlinearity; layer count alone is not the crucial ingredient.

  • Treating the hidden layer as useful even when its units use only linear activation functions.

    The stated approximation result depends on nonlinear activation functions, with sigmoid units as one example.

    Fix: Include suitable nonlinear units when reasoning about the universal approximation property.

  • Forgetting the domain conditions in the approximation statement.

    The source states the result for continuous functions on a compact input region.

    Fix: Include both conditions when stating the property.

Practice the distinction

MEDIUM

A learner says: A network with five linear layers must be more expressive than a network with no hidden layers, and the universal approximation property proves that every such network can approximate any continuous target. Identify both errors and replace the statement with a source-grounded version.

Hints
  • Consider what happens when linear functions are composed.
  • Look closely at the words sufficiently large, single hidden layer, sigmoid units, continuous function, and compact input region.
  • Separate a claim that a suitable network exists from a claim about every individual network.

What do you think happens?

A network has several layers, but all activation functions are linear. After the layers are composed, what kind of function does the whole network represent?

  • A linear function
  • A guaranteed approximation of every continuous function
  • A function whose class is richer solely because more layers were added
Reveal answer

Answer: A linear function

Linear functions composed with linear functions remain linear, so the stack is equivalent to a network with no hidden layers for this purpose.

What to remember

  1. A sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.
  2. Nonlinear activation functions are the crucial ingredient behind this approximation ability.
  3. A network with no hidden layers represents only a very small fraction of possible input-output functions.
  4. Several layers with only linear activation functions still compose to a linear function and do not create a richer function class.
  5. The universal approximation property is a capacity claim about the existence of a sufficiently large suitable network, not a guarantee about every individual network.

Key Takeaways

  • The universal approximation property concerns sufficiently large single-hidden-layer feedforward networks with sigmoid units.
  • The target must be continuous and defined on a compact input region in the stated result.
  • Nonlinearity allows hidden layers to expand the class of functions that can be approximated.
  • Stacking linear layers does not help because their composition remains linear.
  • The property describes what some suitable networks can achieve, not what every individual network automatically achieves.