Era 3 · The deep learning explosion · 2014

25 VGG

Very Deep Convolutional Networks for Large-Scale Image Recognition · Simonyan & Zisserman · ICLR 2015
🟦 this write-up is enough~20–30 minoriginal ↗
The gist in 20 seconds. A very deep CNN with an extremely plain design: nothing but 3×3 convolutions and 2×2 pooling, stacked 16–19 layers deep. The idea: two 3×3 layers give a 5×5 receptive field with fewer parameters and more nonlinearity → depth matters more than filter size.

Context

After AlexNet (#16) comes the CNN architecture race. Simonyan and Zisserman (Oxford) test what pure DEPTH does, under an extremely simple, uniform design.

The idea and the mechanism

Use only small 3×3 convolutions (stride 1) and 2×2 max-pooling, stack them 16–19 layers deep and double the channel count after every pooling step. That uniformity of design is VGG's defining feature.

linear algebra Why a stack of 3×3 beats one big filter

Receptive field. Two 3×3 layers in a row "see" a 5×5 region, three see 7×7: each further layer widens the reach by 2.

Parameters. A 3×3 convolution with C channels in and out costs 3·3·C² = 9C². Compare at equal receptive field:

two 3×3: 2·9C² = 18C²  vs  one 5×5: 25C²
three 3×3: 27C²  vs  one 7×7: 49C²

A stack of small filters gives the same receptive field more cheaply in parameters and with more nonlinearities (there is a ReLU between the convolutions). Depth with small filters wins on both counts — that is VGG's central lesson.

PyTorch A VGG block
import torch.nn as nn
# two 3×3 (receptive field of a 5×5) + ReLU + pooling
block = nn.Sequential(
    nn.Conv2d(C, C, 3, padding=1), nn.ReLU(inplace=True),
    nn.Conv2d(C, C, 3, padding=1), nn.ReLU(inplace=True),
    nn.MaxPool2d(2),
)
3×3 one 3×3 sees 3×3, two 3×3 in a row see 5×5 …cheaper and more nonlinear than one 5×5
A stack of small convolutions widens the receptive field step by step, while staying cheaper and "more nonlinear" than one big filter.
Analogy. To take in a wide panorama you do not need one giant lens — you can walk a few steps back, taking a shot at each. Every "step" (a 3×3 layer) widens the view a little, and between steps you get to think (the ReLU). The total reach is the same as an eye-wateringly expensive wide-angle, but more flexible and cheaper.

Why it matters

VGG-16/19 is a very regular architecture. Heavy though it is (138M parameters), and beaten on classification by GoogLeNet at ILSVRC-2014, its uniformity made it the favourite backbone for feature extraction and perceptual loss (style transfer, quality metrics) for years.

Connections

← builds on16. AlexNet

VGG continues the AlexNet line but systematises it: instead of filters of assorted sizes, uniform 3×3 ones, and more depth. It is AlexNet tidied up, boiled down to the clear principle "depth plus small filters".

→ runs into27. ResNet

VGG pushed depth to 19 layers, but beyond that simply stacking more stopped working (degradation). A year later ResNet breaks through that wall with skip connections and takes depth to 152 — a direct continuation of the question "how do we make networks deeper still".

↔ contrast32. Transformer

VGG is the apotheosis of hard-wired inductive bias (locality of convolution, hierarchy). The Transformer and ViT go the other way — minimal built-in assumptions, everything learned from data. History swings between these two poles.

Questions worth asking

If depth with small filters is so good, why not build networks of 3×3 layers of unlimited depth?

Because simply stacking more runs into an optimization problem: in very deep networks even the training error goes up (degradation, not overfitting). VGG-19 is close to the limit of what trains head-on. ResNet will break the wall with skip connections — without them, depth beyond ~20 layers stops helping.

VGG lost to GoogLeNet — why is it still so well liked?

For its simplicity and uniformity. GoogLeNet, with its inception modules, is more complicated and more temperamental as a backbone. VGG is predictable: an even stack of 3×3 layers is easy to cut into features of different levels, which is why it became the standard for transfer learning, perceptual loss and style transfer — a win in convenience rather than in a competition metric.

Where do 138M parameters come from if all the filters are small?

Almost all of them sit in the fully connected layers at the end (fc-4096), not in the convolutions. That was fixed later: global average pooling in place of the big FC layers cuts parameters by an order of magnitude at the same quality. VGG is a clear illustration that the "weight" of a network is often not where you think.

What to read in the original

The idea is simple — this write-up is enough. The main thing to take away is the "depth plus small filters" principle and why a stack of 3×3 is more efficient (the math box).