Concepts / Expressive Power of Neural Networks

Expressive Power of Neural Networks

Neural networks with sign activation can be understood through geometric illustrations.

  • Programming

From Computation to Geometry

A neural network is often described as a sequence of numerical transformations. With sign activation, there is another useful viewpoint: the computation can be understood as dividing the input space into regions and assigning outputs from the set {−1, 1}. For two-dimensional real inputs, these regions can be illustrated geometrically. The key is to follow two stages: hidden neurons first perform geometric tests, and the output neuron then combines their binary results.

apply sign testone sideother sideInput spacetwo-dimensional real inputsHyperplanegeometric boundaryPositive regionoutput 1Negative regionoutput −1
What regions of geometric space correspond to the positive and negative outputs of a sign-activated neuron?

Hidden Neurons as Halfspaces

Each neuron in the hidden layer of a depth 2 network implements a halfspace predictor. Geometrically, its computation uses a hyperplane as a boundary. The input space is divided into two sides of that boundary, and the sign activation reports a binary result for the side containing the input. Thus, a hidden neuron is not merely producing an intermediate number: it is testing whether an input lies in one of two geometric regions.

hidden-neuron testone sideother sideInput spacebefore the testHyperplaneboundaryHalfspaceselected sideHalfspaceother side
How does each hidden neuron use a hyperplane to divide the input space and identify one side of it?

The word halfspace is important because the hidden neuron’s decision is geometric rather than arbitrary. One hyperplane supplies the boundary, and the two sides of that boundary supply the possible regions recognized by the sign-activated neuron. With several hidden neurons, the network has several such tests available for the same input.

Combining Hidden Decisions

The output layer provides the second stage of the geometric interpretation. The single output neuron applies a halfspace to the binary outputs produced by the hidden neurons. In this representation, the output neuron does not inspect the original input space directly. It receives the hidden neurons’ binary decisions and applies another halfspace test to that binary-output space.

binary resultbinary resultbinary resultclassifyHidden test 1binary outputOutput halfspacecombines binary resultsNetwork outputinside or outsideHidden test 2binary outputHidden test kbinary output
How does the output neuron combine the hidden neurons' halfspace decisions to classify points inside or outside the resulting region?

A halfspace can implement conjunction, so the output stage can require multiple hidden-layer tests to agree. In the input space, this combination is interpreted as an intersection of halfspaces. The resulting accepted region can therefore be a convex polytope rather than just the region from one individual hidden neuron.

Depth 2 as a Polytope Builder

Counting Faces from Hidden Neurons

A depth 2 network has k hidden neurons and represents a convex polytope under the stated geometric interpretation. How many faces does the polytope have?

Identify the hidden tests: Each hidden neuron implements a halfspace predictor, so the hidden layer supplies k geometric constraints.

Interpret the output stage: The output neuron applies a halfspace to the hidden neurons' binary outputs. This can combine the hidden tests as an intersection of halfspaces.

Use the network-to-polytope relationship: For k hidden neurons, the resulting network can express a convex polytope with k − 1 faces.

The convex polytope represented by the depth 2 network has k − 1 faces.

testtesttestcombinecombinecombinerepresentInput pointpoint in input spaceHalfspace 1hidden testIntersectioncombined regionConvex polytopek − 1 facesHalfspace 2hidden testHalfspace khidden test
How do the hidden-layer halfspaces combine to form the boundary and interior of a convex polytope?

What do you think happens?

A depth 2 network has 6 hidden neurons and follows the stated relationship between hidden neurons and polytope faces. How many faces does the represented convex polytope have?

  • 5
  • 6
  • 7
Reveal answer

Answer: 5

The relationship is k − 1 faces for k hidden neurons. Substituting k = 6 gives 6 − 1 = 5 faces.

contributes to boundarycontributes to boundarycontributes to boundaryHidden neuron 1halfspace constraintk − 1 facesconvex polytopeHidden neuron 2halfspace constraintHidden neuron khalfspace constraint
How does each relevant hidden-layer constraint correspond to a face of the convex polytope, and how can the total number of faces be counted?

Changing Hidden-Layer Capacity

Network descriptionGeometric interpretation
One hidden neuronOne halfspace predictor
Several hidden neuronsSeveral halfspace tests whose results can be combined
Depth 2 network with k hidden neuronsAn intersection-of-halfspaces representation that can express a convex polytope with k − 1 faces

Increasing the number of hidden neurons increases the number of halfspace predictors available to the network. The hidden layer can therefore supply more geometric tests for the output stage to combine. In the specific depth 2 relationship described here, k hidden neurons correspond to a convex polytope with k − 1 faces. This gives a direct way to connect a network-size question with a geometric face-count question.

When analyzing a sign-activated network geometrically, separate the two stages. First identify the halfspace test performed by each hidden neuron. Then examine how the output neuron combines the resulting binary values. This prevents the common mistake of treating the entire network as one unexplained boundary.

Common Interpretation Errors

  • Treating a hidden neuron as the complete polytope.

    A hidden neuron implements one halfspace predictor. The final geometric region comes from combining hidden-layer results through the output neuron.

    Fix: Interpret each hidden neuron as one geometric test, then analyze their intersection as produced by the output stage.

  • Ignoring the binary nature of sign-activated hidden outputs.

    The geometric interpretation in the source uses binary outputs in {−1, 1}.

    Fix: Track each hidden neuron as returning one of the two sign outputs associated with the two sides of its boundary.

  • Counting k faces for k hidden neurons.

    The stated relationship for the described depth 2 network is k − 1 faces.

    Fix: Substitute the number of hidden neurons into k − 1. For 6 hidden neurons, the count is 5 faces.

  • Describing the output neuron as operating directly on the original input coordinates.

    The output neuron applies a halfspace to the binary outputs produced by the hidden neurons.

    Fix: Explain the output neuron as the second-stage combiner of hidden-layer decisions.

Check Your Understanding

MEDIUM

A depth 2 network uses sign activation and has 9 hidden neurons. Explain the geometric role of one hidden neuron, describe what the output neuron does with the hidden neurons’ binary results, and determine the number of faces in the convex polytope represented by the network under the stated relationship.

Hints
  • Begin by identifying the hyperplane and the two halfspaces associated with one hidden neuron.
  • The output neuron applies a halfspace to the hidden neurons’ binary outputs and can implement their conjunction.
  • Use k − 1 with k = 9.
  1. A sign-activated neural network can be read as a geometric construction. Each hidden-layer neuron is a halfspace predictor: a hyperplane divides the input space, and the sign output identifies one of the two sides. The output neuron applies a halfspace to the hidden neurons’ binary outputs, allowing their tests to be combined as an intersection of halfspaces. For the described depth 2 network, k hidden neurons can represent a convex polytope with k − 1 faces.

Key Takeaways

  • Sign activation allows a neural network to be interpreted through regions and shapes in the input space.
  • Each hidden-layer neuron in the depth 2 network acts as a halfspace predictor.
  • The output neuron applies a halfspace to the hidden neurons’ binary outputs and can combine their tests as an intersection.
  • For k hidden neurons, the described network can express a convex polytope with k − 1 faces.