Era 2 · Foundations · 1989 → 1998

9 LeNet / CNN

Gradient-Based Learning Applied to Document Recognition · LeCun, Bottou, Bengio, Haffner · Proc. IEEE
🟧 read selectively~2–3 horiginal ↗
The gist in 20 seconds. A convolutional network built on three ideas — local receptive fields, weight sharing (filters) and pooling — learns features straight from pixels, invariant to shift and with few parameters. The template for every modern CNN.

Context

A fully connected network on an image is a disaster: millions of weights and total disregard for structure (neighbouring pixels are related, an object can move). Yann LeCun (1989→1998) builds a network with that structure wired into the architecture.

The idea and the mechanism

Three principles. Local receptive fields: a neuron looks at a small patch. Weight sharing: one set of weights (a filter) slides across the whole image — this is the convolution operation:

(I ∗ K)(i, j) = Σm Σn I(i+m, j+n) · K(m, n)

One filter = a feature detector (an edge, a corner) that does not care where the feature sits. Pooling aggregates neighbouring responses → robustness to shifts and a drop in dimensionality. A stack of conv → pool → conv → pool → fc, trained end-to-end by backprop; early layers catch edges, later ones parts of objects.

linear algebra Weight sharing: how many parameters convolution saves

Take a 32×32 input and a hidden layer of 32×32 neurons. The fully connected version: each of the 1024 outputs is connected to each of the 1024 inputs →

1024 × 1024 ≈ 106 parameters

The convolutional version with a 5×5 filter: the same filter is applied at every position →

5 × 5 = 25 parameters  (per channel, whatever the image size)

A saving of tens of thousands of times, and the same thing gives you translational equivariance: move the object → the response moves the same way, because the filter is identical everywhere. Convolution encodes the prior knowledge "locality + invariance to position" directly into the structure — the strongest inductive bias there is for images.

NumPy Convolving one filter over an image
import numpy as np

def conv2d(img, K):                # img: (H,W), filter K: (k,k)
    k = K.shape[0]
    H, W = img.shape
    out = np.zeros((H - k + 1, W - k + 1))
    for i in range(out.shape[0]):
        for j in range(out.shape[1]):
            patch = img[i:i+k, j:j+k]
            out[i, j] = (patch * K).sum()   # one filter at every position
    return out                              # feature map
image conv pool conv pool fc class edges → parts → objects → class
The CNN pipeline: convolutions (which learn features) alternating with pooling (robustness + compression), then a fully connected layer → class.
Analogy. A filter is a stencil stamp: you press the same stamp — "is there a vertical edge here?" — all over the picture. There is no need to learn a separate detector for every corner of the image — one stamp looks for the feature everywhere. Pooling then reports "the stamp fired somewhere around here", blurring the exact location.

Why it matters

LeNet-5 read handwritten digits in production (cheques, postcodes) and became the template for every CNN. The first convincing demonstration that a learned hierarchy of features from pixels beats hand-crafted ones. The idea will reach full force in 2012 (AlexNet, #16), once GPUs and ImageNet arrive.

Connections

← builds on7. Backpropagation

A CNN is the same network, trained by backprop, but with constraints wired in (locality, sharing). A convolution is simply a layer with shared weights; the gradient flows through it by the same chain rule.

→ leads to16. AlexNet

AlexNet is LeNet grown up on GPUs and ImageNet: the same convolutions + pooling, but deeper, with ReLU and dropout. A 1998 idea waited for the data and the hardware — and then took off.

ViT drops the hard-wired inductive bias of convolutions and replaces it with data scale + self-attention. A CNN "knows" about locality from birth; a ViT "learns" it, but only on large data. The argument "built-in biases vs learn it from data" runs right through ML.

Questions worth asking

Convolution gives translational equivariance, not invariance — what is the difference and does it matter?

Equivariance: move the input → the output moves the same way (a property of convolution itself). Invariance: move the input → the output does not change (what you need for classification, "it's a cat wherever it is"). Invariance is picked up by pooling and by global aggregation at the end. Confusing the two is a common mistake; a CNN is equivariant by construction and invariant only after pooling.

If convolution is so good, why wasn't it invented and adopted widely before 2012?

It was there (LeNet was reading cheques in the 1990s), but it ran into data and compute: small datasets + weak CPUs gave deep CNNs no room, while SVMs were competitive on those tasks and simpler. ImageNet (#15) and GPUs removed both limits — and CNNs immediately pulled ahead.

Pooling throws away information about exact position — isn't that harmful?

It is a deliberate trade-off: we sacrifice precise localization for robustness to shifts and for compression. For classifying "what is in the picture" that is useful. But for tasks where position matters (segmentation, detection) aggressive pooling gets in the way — there people use strided convolutions, dilated convolutions, or drop pooling altogether to keep spatial precision.

What to read in the original

The 1998 paper is long (46 pp.) — read selectively, the sections on convolution, pooling and weight sharing; the graph-transformer part at the end can be skipped.