Dropout at Training vs Inference
Try it: Dropout at Training vs Inference
How dropout zeroes a random subset of a layer's outputs during training (with a visible seeded mask), how inverted dropout scales the survivors by 1 / (1 - rate) while inference is a no-op (or classic dropout scales at inference instead), and why the average training output matches the inference output.
How it works
- For each unit draw u from the seeded generator; keep it when u >= rate, otherwise output 0.
- Inverted dropout (Keras): multiply kept units by 1 / (1 - rate). Classic dropout: leave them unchanged.
- At inference nothing is dropped: inverted dropout is a no-op, classic dropout multiplies by (1 - rate).
- Average the training output over more and more masks: it converges to the inference output.
Default run (18 steps): A layer outputs 6 activations [0.8, 1.5, 0.2, 2, 1.1, 0.6]. Dropout rate 0.5, training mode, inverted dropout (as Keras does). … Average over 1000 masks (seeds 1..1000): [0.7936, 1.56, 0.198, 1.96, 1.1242, 0.5928]; largest gap to the inference output [0.8, 1.5, 0.2, 2, 1.1, 0.6] is 0.06. The average training output converges to the inference output: dropout preserves the expected activation.
Simplified: One activation vector with up to 10 units and a seeded teaching PRNG (mulberry32) instead of TensorFlow's generator.
Loading the simulation…