25 VGG
Context
After AlexNet (#16) comes the CNN architecture race. Simonyan and Zisserman (Oxford) test what pure DEPTH does, under an extremely simple, uniform design.
The idea and the mechanism
Use only small 3×3 convolutions (stride 1) and 2×2 max-pooling, stack them 16–19 layers deep and double the channel count after every pooling step. That uniformity of design is VGG's defining feature.
linear algebra Why a stack of 3×3 beats one big filter
Receptive field. Two 3×3 layers in a row "see" a 5×5 region, three see 7×7: each further layer widens the reach by 2.
Parameters. A 3×3 convolution with C channels in and out costs 3·3·C² = 9C². Compare at equal receptive field:
A stack of small filters gives the same receptive field more cheaply in parameters and with more nonlinearities (there is a ReLU between the convolutions). Depth with small filters wins on both counts — that is VGG's central lesson.
PyTorch A VGG block
import torch.nn as nn
# two 3×3 (receptive field of a 5×5) + ReLU + pooling
block = nn.Sequential(
nn.Conv2d(C, C, 3, padding=1), nn.ReLU(inplace=True),
nn.Conv2d(C, C, 3, padding=1), nn.ReLU(inplace=True),
nn.MaxPool2d(2),
)
Why it matters
VGG-16/19 is a very regular architecture. Heavy though it is (138M parameters), and beaten on classification by GoogLeNet at ILSVRC-2014, its uniformity made it the favourite backbone for feature extraction and perceptual loss (style transfer, quality metrics) for years.
Connections
VGG continues the AlexNet line but systematises it: instead of filters of assorted sizes, uniform 3×3 ones, and more depth. It is AlexNet tidied up, boiled down to the clear principle "depth plus small filters".
VGG pushed depth to 19 layers, but beyond that simply stacking more stopped working (degradation). A year later ResNet breaks through that wall with skip connections and takes depth to 152 — a direct continuation of the question "how do we make networks deeper still".
VGG is the apotheosis of hard-wired inductive bias (locality of convolution, hierarchy). The Transformer and ViT go the other way — minimal built-in assumptions, everything learned from data. History swings between these two poles.
Questions worth asking
If depth with small filters is so good, why not build networks of 3×3 layers of unlimited depth?
Because simply stacking more runs into an optimization problem: in very deep networks even the training error goes up (degradation, not overfitting). VGG-19 is close to the limit of what trains head-on. ResNet will break the wall with skip connections — without them, depth beyond ~20 layers stops helping.
VGG lost to GoogLeNet — why is it still so well liked?
For its simplicity and uniformity. GoogLeNet, with its inception modules, is more complicated and more temperamental as a backbone. VGG is predictable: an even stack of 3×3 layers is easy to cut into features of different levels, which is why it became the standard for transfer learning, perceptual loss and style transfer — a win in convenience rather than in a competition metric.
Where do 138M parameters come from if all the filters are small?
Almost all of them sit in the fully connected layers at the end (fc-4096), not in the convolutions. That was fixed later: global average pooling in place of the big FC layers cuts parameters by an order of magnitude at the same quality. VGG is a clear illustration that the "weight" of a network is often not where you think.
What to read in the original
The idea is simple — this write-up is enough. The main thing to take away is the "depth plus small filters" principle and why a stack of 3×3 is more efficient (the math box).