19 Dropout
Context
Large networks overfit — they memorize noise and generalize badly. Srivastava, Hinton et al. offer a simple, general regularizer.
The idea and the mechanism
At every training step each neuron is switched off (zeroed) with probability 1−p. The network cannot rely on particular neurons and their joint co-adaptation → it is forced to learn redundant, robust features. At inference all neurons are active and the outputs are rescaled so the expected magnitude is preserved.
probability Dropout as ensemble averaging
Take a neuron with activation a that is kept with probability p: the mask is m ∼ Bernoulli(p), the output y = m · a. Then the expected output is:
So at test time, when every neuron is on, the output is multiplied by p — to make the average magnitude match training (or you use "inverted dropout": divide by p during training and leave test time alone).
The main interpretation. Each step trains one of 2n "thinned" subnetworks (one per subset of the n neurons), and they all share weights. Test-time rescaling approximately computes the geometric mean of the predictions of that entire exponential ensemble. In other words, dropout ≈ training a giant ensemble at the cost of a single network.
NumPy Inverted dropout (train / eval)
import numpy as np
def dropout(a, p=0.5, train=True):
if not train:
return a # at inference — unchanged
mask = (np.random.rand(*a.shape) < p) / p # mask + 1/p scaling
return a * mask # some neurons are zeroed
Why it matters
Cheap, general, and for a long time the standard in fully connected and recurrent networks; conceptually it ties regularization to ensembling. In modern CNNs it has partly been displaced by BatchNorm and by sheer volumes of data, but in Transformers (dropout in attention/FFN) it lives on.
Connections
Dropout in the fully connected layers was one of the key tricks that kept the 60M-parameter AlexNet from overfitting ImageNet. The regularizer and the 2012 breakthrough came out of the same lab and worked together.
Both stabilize/regularize training, but differently: dropout injects noise by switching neurons off, BatchNorm normalizes activations (and throws in mild regularization as a side effect). With the arrival of BatchNorm and large datasets, dropout was pushed well back in convolutional networks.
In both cases the strength comes from an ensemble of diverse models. A forest builds many separate trees; dropout hides an exponential ensemble of subnetworks inside a single network with shared weights. Different routes to the same wisdom of crowds.
Questions worth asking
Why does test-time scaling by p only approximate the ensemble rather than compute it exactly?
Exact averaging would mean running all 2n subnetworks and averaging them — impossible. Multiplying by p is a careful approximation to the geometric mean of their outputs, exact for linear networks and good for moderate nonlinearities. In practice the approximation works very well, even though it is not strictly equal to the ensemble.
Why has dropout all but disappeared from modern CNNs, yet stayed in Transformers?
In CNNs its role as a regularizer was largely taken over by BatchNorm, augmentation and abundant data, and dropout between convolutional maps interfered with BN's statistics. Transformers have no BatchNorm (they use LayerNorm), the models are huge and prone to overfitting particular data — so dropout in attention and the FFN remains useful. The choice of regularization depends on the architecture.
Does switching neurons off at random every step hurt convergence?
It slows it down — training is noisier and needs more epochs (you are effectively training many subnetworks at once). It is a deliberate trade: slightly slower training for markedly better generalization. So dropout is used where overfitting is a real threat (big networks, little data), and dropped when data is plentiful.
What to read in the original
The write-up plus the key passages is enough: the masking idea, the test-time scaling and the ensemble interpretation. The long empirical section can be skimmed.