Activation Functions and Their Gradients
Interactive lab
Try it: Activation Functions and Their Gradients
How ReLU, sigmoid and tanh transform a neuron's pre-activation z = w*x + b, and how each function's derivative (the local gradient used by backpropagation) shrinks towards zero in its saturation zones.
How it works
- Form the pre-activation z = w*x + b (or set z directly).
- Apply ReLU max(0, z), sigmoid 1 / (1 + e^-z) and tanh(z).
- Evaluate each derivative: ReLU 1 for z > 0 else 0, sigmoid s(1 - s), tanh 1 - tanh(z)^2.
- Flag saturation: a derivative below 0.05 (or ReLU's zero slope for z <= 0) passes almost no gradient backwards.
Default run (6 steps): A neuron with weight w = 1.5, input x = 2 and bias b = -1. Next: form its pre-activation z = w*x + b. … At z = 2: relu 2 (slope 1), sigmoid 0.8808 (slope 0.105), tanh 0.964 (slope 0.0707). Every function still passes a useful gradient here.
Simplified: One scalar neuron, not a full layer. ReLU's derivative at exactly z = 0 is taken as 0 (the TensorFlow convention). The 0.05 saturation threshold is a teaching choice.
Educational simulation
Loading the simulation…