Activation Functions
A feedforward neural network is a directed acyclic graph whose nodes are neurons and whose directed edges are weighted connections.
Try it: Neural Network Forward Pass
How a neural network turns inputs into an output: each neuron computes a weighted sum plus bias, applies an activation function, and passes the result on.
How it works
- Each hidden neuron computes z = w1·x1 + w2·x2 + b.
- It applies an activation: ReLU keeps positive values and zeroes negatives; sigmoid squashes z into 0–1.
- The output neuron does the same with the hidden activations as its inputs.
- Changing any input, weight or bias changes every value downstream of it.
Default run (7 steps): Inputs x1 = 1, x2 = 0.5. Activation: ReLU. … Prediction y = ReLU(1.25) = max(0, 1.25) = 1.25.
Simplified: Tiny 2-2-1 network with hand-set weights, not a trained model. No training happens here. Activations offered: ReLU, sigmoid and sign (texts differ on the value of sign at exactly 0; here it is 0).
Loading the simulation…
A One-Way Computation
A feedforward neural network is a one-way route for information. Its nodes are neurons, and its directed edges are weighted connections. Values enter through the input layer, move through connected neurons, and eventually reach the output layer. Because the connections are directed and the network is acyclic, information does not travel backward or form a cycle.
Layers and Their Roles
The full neuron set is divided into layer sets V0 through VT. The input layer is V0, hidden layers lie between the input and output layers, and the output layer is VT. The layer index records the processing order: values begin in V0, pass through hidden layers when they exist, and finish in VT.
These names describe roles in the computation. The input layer exposes starting values, hidden layers perform intermediate transformations, and the output layer contains the network's final output. They are not names for unrelated collections of neurons; they identify positions in one directed flow.
Features and the Constant Neuron
For an n-dimensional input space, the described input layer contains n feature neurons plus one constant neuron. Each feature neuron represents one input feature. The additional constant neuron always outputs 1, so it is part of the input representation even though it is not a changing feature value.
Inside a Neuron
A neuron does not pass the combined influence of its inputs directly to its output. First, it receives a scalar input formed from a weighted sum of the outputs of connected predecessor neurons. It then applies an activation function to that single scalar. The activation function is the rule that converts the neuron's combined input into its output.
Tracing One Neuron
A neuron receives outputs from connected predecessor neurons. Describe the order of operations without choosing a particular activation function.
Receive: The neuron receives outputs from its connected predecessor neurons.
Combine: The incoming outputs are combined using their weighted connections to form one scalar input.
Activate: The neuron applies its chosen activation function to that scalar input.
Output: The result of the activation function becomes the neuron's output and can supply a later connected neuron.
The activation function is applied after the weighted sum and before the neuron passes its output onward.
Tracing a Layered Route
Consider a route with an input layer, a hidden layer, and an output layer. The input neurons expose feature values and the constant value. A hidden neuron receives weighted outputs from connected input neurons, forms its scalar input, and applies an activation function. The output neuron then receives weighted outputs from connected hidden neurons, forms its own scalar input, and applies an activation function before producing the network's final output.
Three Activation Behaviors
| Function | How it transforms the scalar input | Main characteristic |
|---|---|---|
| Sign | Bases the output on the sign of the scalar input | Sign-based transformation |
| Threshold | Uses a cutoff-based transformation | Abrupt change around a cutoff |
| Sigmoid | Uses a smooth formula that approximates threshold behavior | Gradual transition rather than an abrupt one |
For the sign function, begin with the weighted sum, call it a, and apply the sign-based rule: σ(a) = sign(a). The threshold function also makes a cutoff-based transformation. The sigmoid function is described as a smooth approximation to threshold behavior because its transition is gradual rather than an abrupt cutoff.
Abrupt and Smooth Transitions
The key contrast between threshold and sigmoid behavior is the transition around the cutoff. A threshold function changes abruptly according to its cutoff rule. A sigmoid function changes smoothly, so it approximates the same general threshold behavior without making the transition abrupt.
Check Your Trace
A network has an input layer, one hidden layer, and an output layer. Explain the order in which a value is processed. Then state where the activation function is applied in each neuron. Finally, classify each description as sign-based, cutoff-based, or smooth approximation to cutoff behavior: sign, threshold, sigmoid.
Hints
- Start with the input layer and follow directed connections forward.
- Separate the weighted-sum step from the activation step.
- Use the defining behavior of each function rather than assuming a numerical convention not supplied in the problem.
Practice Answer
Trace the processing route and identify the role of each activation function.
Route: Values begin at the input layer, move through weighted connections to hidden neurons, and then move through weighted connections to the output layer.
Neuron operation: Each receiving neuron first forms a scalar weighted sum from connected predecessor outputs and then applies its activation function.
Function classification: Sign is sign-based, threshold is cutoff-based, and sigmoid is a smooth approximation to threshold behavior.
The activation function is the second stage inside the neuron: it converts the weighted sum into the neuron's output.
Common Mistakes
Treating a feedforward network as a collection of independent neurons
The neurons and weighted communication links form a directed graph, and the links determine how one stage supplies information to the next.
Fix:
Trace the route from the input layer through hidden layers to the output layer.Calling the weighted sum the final neuron output
The weighted sum is the scalar input to the activation function, not necessarily the neuron's final output.
Fix:
Apply the activation function after forming the weighted sum.Putting the constant neuron in a hidden layer
For an n-dimensional input space, the described input layer contains the n feature neurons plus the constant neuron.
Fix:
Keep the constant neuron in the input layer and remember that it always outputs 1.Treating sigmoid and threshold as identical
Threshold behavior is cutoff-based, while sigmoid behavior is smooth and approximates threshold behavior.
Fix:
Look for the distinction between an abrupt change and a gradual transition.Assuming a numerical threshold convention that was not supplied
The supplied material describes threshold behavior as cutoff-based but does not specify a complete numerical convention.
Fix:
Use the exact definition given in the exercise or lesson.
Key Takeaways
- A feedforward neural network is a directed acyclic graph of neurons connected by weighted directed edges.
- The input layer is V0, hidden layers lie between the input and output layers, and the output layer is VT.
- An input layer for an n-dimensional input includes n feature neurons and one constant neuron that always outputs 1.
- A neuron first receives a weighted sum from predecessor outputs and then applies an activation function to that scalar.
- Sign is sign-based, threshold is cutoff-based, and sigmoid provides a smooth approximation to threshold behavior.
Key Takeaways
- Feedforward networks move information forward through a directed acyclic graph.
- Input, hidden, and output layers identify positions and roles in that flow.
- A neuron applies its activation function after calculating a weighted sum.
- Sign, threshold, and sigmoid functions differ in whether their transformation is sign-based, cutoff-based, or smooth.
- Sigmoid is called a smooth approximation to threshold because it replaces an abrupt transition with a gradual one.