Concepts / Activation Functions

Activation Functions

A feedforward neural network is a directed acyclic graph whose nodes are neurons and whose directed edges are weighted connections.

  • Programming
Interactive lab

Try it: Neural Network Forward Pass

How a neural network turns inputs into an output: each neuron computes a weighted sum plus bias, applies an activation function, and passes the result on.

How it works

  1. Each hidden neuron computes z = w1·x1 + w2·x2 + b.
  2. It applies an activation: ReLU keeps positive values and zeroes negatives; sigmoid squashes z into 0–1.
  3. The output neuron does the same with the hidden activations as its inputs.
  4. Changing any input, weight or bias changes every value downstream of it.

Default run (7 steps): Inputs x1 = 1, x2 = 0.5. Activation: ReLU. … Prediction y = ReLU(1.25) = max(0, 1.25) = 1.25.

Simplified: Tiny 2-2-1 network with hand-set weights, not a trained model. No training happens here. Activations offered: ReLU, sigmoid and sign (texts differ on the value of sign at exactly 0; here it is 0).

Educational simulation

Loading the simulation…

A One-Way Computation

A feedforward neural network is a one-way route for information. Its nodes are neurons, and its directed edges are weighted connections. Values enter through the input layer, move through connected neurons, and eventually reach the output layer. Because the connections are directed and the network is acyclic, information does not travel backward or form a cycle.

weighted connectionweighted connectionweighted connectionweighted connectionInput featuresstarting valuesHidden neuronsintermediate transformationOutput neuronsfinal outputHidden neuronsintermediate transformation
How does information move from input features through hidden neurons to the output without flowing backward or forming a cycle?

Layers and Their Roles

The full neuron set is divided into layer sets V0 through VT. The input layer is V0, hidden layers lie between the input and output layers, and the output layer is VT. The layer index records the processing order: values begin in V0, pass through hidden layers when they exist, and finish in VT.

containscontainscontainsNeuron set Vfull networkInput layer V0starting valuesHidden layersintermediatetransformationsOutput layer VTfinal output
What contains what in a layered network, and how are the input, hidden, and output layers different?

These names describe roles in the computation. The input layer exposes starting values, hidden layers perform intermediate transformations, and the output layer contains the network's final output. They are not names for unrelated collections of neurons; they identify positions in one directed flow.

Features and the Constant Neuron

For an n-dimensional input space, the described input layer contains n feature neurons plus one constant neuron. Each feature neuron represents one input feature. The additional constant neuron always outputs 1, so it is part of the input representation even though it is not a changing feature value.

containscontainscontainsInput layern features plus oneconstantFeature neuronfeature valueFeature neuronfeature valueConstant neuronalways outputs 1
How are input features represented as neurons, and where does the constant neuron fit into the input representation?

Inside a Neuron

A neuron does not pass the combined influence of its inputs directly to its output. First, it receives a scalar input formed from a weighted sum of the outputs of connected predecessor neurons. It then applies an activation function to that single scalar. The activation function is the rule that converts the neuron's combined input into its output.

combine with weightssingle scalarproducePredecessor outputsconnected neuronsWeighted sumscalar inputActivation functionconversion ruleNeuron outputresult
What happens to the weighted sum received from predecessor neurons before the neuron produces its output?

Tracing One Neuron

A neuron receives outputs from connected predecessor neurons. Describe the order of operations without choosing a particular activation function.

Receive: The neuron receives outputs from its connected predecessor neurons.

Combine: The incoming outputs are combined using their weighted connections to form one scalar input.

Activate: The neuron applies its chosen activation function to that scalar input.

Output: The result of the activation function becomes the neuron's output and can supply a later connected neuron.

The activation function is applied after the weighted sum and before the neuron passes its output onward.

Tracing a Layered Route

Consider a route with an input layer, a hidden layer, and an output layer. The input neurons expose feature values and the constant value. A hidden neuron receives weighted outputs from connected input neurons, forms its scalar input, and applies an activation function. The output neuron then receives weighted outputs from connected hidden neurons, forms its own scalar input, and applies an activation function before producing the network's final output.

weighted connectionsactivationweighted connectionsactivationInput featuresfeature neurons andconstant neuronHidden weighted sumscalar inputOutput weighted sumscalar inputHidden outputafter activationFinal outputafter activation
How does a value move through a small network from feature representation to final output?

Three Activation Behaviors

FunctionHow it transforms the scalar inputMain characteristic
SignBases the output on the sign of the scalar inputSign-based transformation
ThresholdUses a cutoff-based transformationAbrupt change around a cutoff
SigmoidUses a smooth formula that approximates threshold behaviorGradual transition rather than an abrupt one
applyapplyapplyScalar inputweighted sumSignsign-basedThresholdcutoff-basedSigmoidsmooth approximation
How do the sign, threshold, and sigmoid functions differ in the kind of output transformation they apply to the same scalar input?

For the sign function, begin with the weighted sum, call it a, and apply the sign-based rule: σ(a) = sign(a). The threshold function also makes a cutoff-based transformation. The sigmoid function is described as a smooth approximation to threshold behavior because its transition is gradual rather than an abrupt cutoff.

Abrupt and Smooth Transitions

The key contrast between threshold and sigmoid behavior is the transition around the cutoff. A threshold function changes abruptly according to its cutoff rule. A sigmoid function changes smoothly, so it approximates the same general threshold behavior without making the transition abrupt.

approacheschanges at cutoffchanges graduallycontinues smoothlyBelow cutoffthreshold input regionLower input regionsigmoid responseCutoffabrupt transitionGradual transitionsmooth changeAbove cutoffthreshold output regionHigher input regionsigmoid response
How does the sigmoid's gradual transition compare with the threshold function's abrupt change around its cutoff?

Check Your Trace

MEDIUM

A network has an input layer, one hidden layer, and an output layer. Explain the order in which a value is processed. Then state where the activation function is applied in each neuron. Finally, classify each description as sign-based, cutoff-based, or smooth approximation to cutoff behavior: sign, threshold, sigmoid.

Hints
  • Start with the input layer and follow directed connections forward.
  • Separate the weighted-sum step from the activation step.
  • Use the defining behavior of each function rather than assuming a numerical convention not supplied in the problem.

Practice Answer

Trace the processing route and identify the role of each activation function.

Route: Values begin at the input layer, move through weighted connections to hidden neurons, and then move through weighted connections to the output layer.

Neuron operation: Each receiving neuron first forms a scalar weighted sum from connected predecessor outputs and then applies its activation function.

Function classification: Sign is sign-based, threshold is cutoff-based, and sigmoid is a smooth approximation to threshold behavior.

The activation function is the second stage inside the neuron: it converts the weighted sum into the neuron's output.

Common Mistakes

  • Treating a feedforward network as a collection of independent neurons

    The neurons and weighted communication links form a directed graph, and the links determine how one stage supplies information to the next.

    Fix: Trace the route from the input layer through hidden layers to the output layer.

  • Calling the weighted sum the final neuron output

    The weighted sum is the scalar input to the activation function, not necessarily the neuron's final output.

    Fix: Apply the activation function after forming the weighted sum.

  • Putting the constant neuron in a hidden layer

    For an n-dimensional input space, the described input layer contains the n feature neurons plus the constant neuron.

    Fix: Keep the constant neuron in the input layer and remember that it always outputs 1.

  • Treating sigmoid and threshold as identical

    Threshold behavior is cutoff-based, while sigmoid behavior is smooth and approximates threshold behavior.

    Fix: Look for the distinction between an abrupt change and a gradual transition.

  • Assuming a numerical threshold convention that was not supplied

    The supplied material describes threshold behavior as cutoff-based but does not specify a complete numerical convention.

    Fix: Use the exact definition given in the exercise or lesson.

Key Takeaways

  1. A feedforward neural network is a directed acyclic graph of neurons connected by weighted directed edges.
  2. The input layer is V0, hidden layers lie between the input and output layers, and the output layer is VT.
  3. An input layer for an n-dimensional input includes n feature neurons and one constant neuron that always outputs 1.
  4. A neuron first receives a weighted sum from predecessor outputs and then applies an activation function to that scalar.
  5. Sign is sign-based, threshold is cutoff-based, and sigmoid provides a smooth approximation to threshold behavior.

Key Takeaways

  • Feedforward networks move information forward through a directed acyclic graph.
  • Input, hidden, and output layers identify positions and roles in that flow.
  • A neuron applies its activation function after calculating a weighted sum.
  • Sign, threshold, and sigmoid functions differ in whether their transformation is sign-based, cutoff-based, or smooth.
  • Sigmoid is called a smooth approximation to threshold because it replaces an abrupt transition with a gradual one.