Concepts / Nonlinear activation functions

Nonlinear activation functions

A sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

  • Programming

Why Layers Alone Are Not Enough

A common first impression is that adding layers automatically makes a neural network capable of representing complex relationships. The important qualification is that hidden layers help only when the units introduce nonlinearity. If every activation function is linear, adding more layers does not create a richer class of functions.

What do you think happens?

What kind of function can a network still represent when several layers are stacked but every activation function is linear?

  • Only a linear function
  • Any continuous function on a compact input region
  • Only a constant function
Reveal answer

Answer: Only a linear function

Linear functions composed with linear functions remain linear. Therefore, a multi-layer network with only linear activations is equivalent in function class to a network with no hidden layers.

The Linear-Stacking Limitation

A unit's activation function determines how it responds to the activation pattern arriving from the network's input units. When that response is linear at every layer, each layer performs a linear transformation of what it receives. Composing one linear function with another still produces a linear function. The same remains true when more linear layers are added. Consequently, the entire multi-layer network represents only one overall linear transformation, rather than a richer nonlinear relationship.

linear mappinglinear mappinglinear mappingequivalent overall mappinglinear outputInputInputLinear layer 1One lineartransformationLinear layer 2OutputOutput
Why does a network with multiple layers of only linear activations still represent just one linear transformation?

Comparing a Linear Stack with a Nonlinear Layer

Consider two feedforward networks. The first has several layers, but every activation function is linear. The second has a hidden layer whose units use a nonlinear activation function. What is the key difference in their representational capacity?

Trace the first network: Each layer applies a linear response to the incoming activation pattern. Because linear functions composed with linear functions remain linear, the whole network is equivalent to one linear transformation.

Trace the second network: The hidden units introduce a nonlinear response between weighted layers. Nonlinearity is the crucial ingredient that allows the hidden layer to expand the class of input-output functions the network can approximate.

Compare the outcomes: The number of layers alone does not determine the richer capability. The decisive difference is whether the activation functions introduce nonlinearity.

A deep stack of only linear activations remains linear, while a hidden layer with suitable nonlinear units can greatly expand the functions the network can approximate.

What Nonlinearity Changes

Nonlinear activation functions change the relationship between the activation pattern entering a unit and the response produced by that unit. This prevents the complete network from collapsing into one overall linear transformation. The hidden layer can therefore contribute behavior that a network with no hidden layers cannot represent. The source specifically identifies sigmoid units as one example of suitable nonlinear units and notes that other nonlinear activation functions can work when they satisfy mild conditions.

arrivesactivatesenablesproducesInput patternWeighted layerNonlinear responsesigmoid unit or anothersuitable nonlinear unitExpanded functionclassOutput
What changes in the input-output mapping when a nonlinear activation is inserted between weighted layers?
direct mappingactivatesmapsInputInputOutputsmall fraction of possiblefunctionsNonlinear hiddenunitssuitable activationfunctionsOutputgreatly expanded functionclass
How does the set of functions representable by a direct linear model differ from the functions representable using a nonlinear hidden layer?

Universal Approximation

The universal approximation property says that a sufficiently large feedforward network with a single hidden layer of sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

The statement has several important parts. It concerns a feedforward network with one hidden layer, not an arbitrary architecture. The hidden layer must be sufficiently large, and the stated example uses sigmoid units. The target function must be continuous, and the input region must be compact. Within those conditions, the approximation can be made as accurate as desired. This is a capacity result: it says that networks of the stated kind have the ability to approximate functions in this way.

definescan produceis approximatedCompact inputregionSufficiently largehidden layersigmoid unitsContinuous targetfunctionApproximationany desired accuracy
How can a sufficiently large single hidden layer combine nonlinear units to approximate an arbitrary continuous function on a compact input region?

Reading the Approximation Claim Carefully

Interpret the statement that a sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.

Identify the architecture: The claim is about a feedforward network with one hidden layer.

Identify the required capacity: The hidden layer must be sufficiently large. The claim does not say that an arbitrarily small hidden layer is enough.

Identify the activation: The stated example uses sigmoid units, and the source notes that other nonlinear activations can work under mild conditions.

Identify the target: The target is any continuous function, considered on a compact input region.

Interpret the result: The network can approximate the target to any desired accuracy. This describes what the architecture can represent in principle; it does not say that every particular network already performs that approximation.

The property is a conditional capacity statement about sufficiently large networks with suitable nonlinear units, continuous targets, and compact input regions.

Capacity Is Not a Guarantee

The universal approximation property should not be read as a claim that every individual network can approximate every function. It says that a sufficiently large network with the stated kind of hidden layer has the capacity to approximate any continuous function on a compact input region to any desired accuracy. A network that is too small, uses unsuitable activations, or does not meet the stated conditions is not covered by that claim.

can approximatedoes not implySufficiently largenetworksone hidden layer withsuitable nonlinear unitsEvery individualnetworkApproximation abilitycontinuous functions oncompact regionsNo universalguarantee
What is the difference between saying that some sufficiently large networks can approximate any target function and saying that every network can do so?
  • Treating the universal approximation property as a guarantee about every network.

    The result requires a sufficiently large hidden layer and applies under stated conditions involving the activation, target function, and input region.

    Fix: Read the result as a capacity claim: networks of the required kind can have the ability to approximate the target.

  • Assuming that adding layers automatically creates nonlinear representational power.

    Linear functions composed with linear functions remain linear.

    Fix: Check whether the activation functions introduce nonlinearity, not merely whether more layers were added.

  • Ignoring the conditions in the approximation statement.

    Those conditions are part of the stated universal approximation property.

    Fix: State the architecture and the conditions together when explaining the result.

Check Your Understanding

MEDIUM

Explain in your own words why a network with several layers of only linear activation functions is equivalent to a network with no hidden layers in terms of the class of functions it can represent. Then state two conditions included in the universal approximation property.

Hints
  • Start with what happens when two linear functions are composed.
  • Remember that the approximation result refers to a sufficiently large single hidden layer with suitable nonlinear units.
  • Include the type of target function and the type of input region.

What do you think happens?

A network has one hidden layer, but every hidden unit uses a linear activation function. Does the presence of the hidden layer alone give it the universal approximation property described here?

  • Yes, because any hidden layer is sufficient
  • No, because the stated result depends on suitable nonlinear activation functions
  • Yes, because depth is the only requirement
Reveal answer

Answer: No, because the stated result depends on suitable nonlinear activation functions

The source identifies nonlinearity as the crucial ingredient. With only linear activations, even multiple layers remain equivalent to one linear transformation.

Key Takeaways

  1. A sufficiently large single hidden layer with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.
  2. Nonlinear activation functions are the crucial ingredient behind this expanded approximation ability.
  3. A network with no hidden layers represents only a very small fraction of possible input-output functions.
  4. Adding multiple layers does not overcome the limitation when every activation function is linear, because compositions of linear functions remain linear.
  5. The universal approximation property describes the capacity of sufficiently large networks under stated conditions, not a guarantee about every individual network.

Key Takeaways

  • Nonlinearity, rather than layer count alone, enables a neural network to approximate a much richer class of functions.
  • A sufficiently large one-hidden-layer feedforward network with sigmoid units can approximate any continuous function on a compact input region to any desired accuracy.
  • Several layers with only linear activations still collapse into one overall linear transformation.
  • Universal approximation is a capacity claim about networks satisfying the stated conditions, not a guarantee about every individual network.