Era 5 · The LLM era · 2020

39 Vision Transformer

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · Dosovitskiy et al. · Google · ICLR 2021
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. A pure Transformer run over a sequence of image patches (no convolutions) beats CNNs — but only given enough pre-training. On mid-sized data it loses: it has none of the built-in inductive biases of convolution, and scale of data is what replaces them.

Context

Vision belonged to CNNs; the Transformer had conquered NLP. Dosovitskiy et al. bring a pure Transformer to vision.

The idea and the mechanism

Cut the image into 16×16 patches, project each one linearly into a vector (like the embedding of a word token), add positional embeddings and a class token, and feed the sequence of patches into a standard Transformer encoder. No convolutions. Self-attention gives you global context from the very first layer.

linear algebra Patches as tokens — and why you need data

An H×W×C image is cut into N = HW/P² patches of size P×P. Each patch is flattened into a vector of length P²C and projected linearly into dimension d; positional embeddings are added:

z0 = [xclass; xp1E; …; xpNE] + Epos

After that it is an ordinary Transformer encoder. The key subtlety: a convolution has locality and translation equivariance hard-wired into it (strong prior assumptions about images). ViT does NOT have them — it has to learn them from data. That is why on mid-sized datasets ViT loses to CNNs and only pulls ahead when pre-trained on very large data (JFT-300M): scale of data substitutes for inductive bias. This is a statement of the fundamental trade-off, "built-in assumptions vs learn it from data".

PyTorch Patchification and linear embedding
import torch
# cut a (C,H,W) image into P×P patches and embed them linearly
def patchify(x, P, E):                 # E: (P*P*C, d)
    C, H, W = x.shape
    p = x.unfold(1, P, P).unfold(2, P, P)   # (C, H/P, W/P, P, P)
    p = p.permute(1, 2, 0, 3, 4).reshape(-1, C * P * P)
    return p @ E                       # (N, d) — patch tokens
16×16 patches tokens Transformer class
The image becomes a sequence of patch tokens and goes into an ordinary Transformer — just like a sentence made of words.
Analogy. Reading a painting not "with a painter's eye" (someone who already knows about edges and shapes — that is the CNN with its built-in rules) but as a text made of tiles, where you have to work out the rules of composition yourself. That is harder and takes far more paintings to "read" (huge data), but it does not constrain the model with someone else's assumptions about how vision works.

Why it matters

It showed the Transformer is universal (one architecture for text and vision) and opened the road to multimodal models (CLIP, multimodal LLMs). And it sharpened the "built-in biases vs scale of data" argument that runs through all of modern ML.

Connections

← builds on32. Transformer

ViT is literally a Transformer encoder fed patches instead of words. No new mechanisms: all the power comes from self-attention, carried over from NLP into vision unchanged.

↔ contrast9. LeNet / CNN

Two poles of one argument. A CNN knows about locality from birth (the convolution), ViT learns it — but only on large data. On small data the CNN's built-in biases win; on large data, ViT's flexibility.

→ leads to40. CLIP

A single Transformer architecture for vision and language is a precondition for multimodality. CLIP ties a visual encoder to a text encoder, and Transformer-based vision (like ViT) is the natural visual backbone for that.

Questions worth asking

If ViT needs a huge dataset, what is its practical point?

Transfer. Pre-train ViT once on gigantic data, then fine-tune it on small target sets — and there it works very well. On top of that, hybrid and improved variants (DeiT with distillation, Swin with local windows) cut the appetite for data by giving some inductive bias back. So "needs data" is about the pre-training, not about every application.

What does ViT "lose" by dropping convolutions?

Guaranteed translation equivariance and efficiency on small data — the things a convolution gives you for free. In exchange it gets global context from the first layer and flexibility without rigid assumptions. This is not "better/worse" but a different trade-off: prior knowledge (CNN) against learning everything from data (ViT).

Why 16×16 patches specifically, rather than one pixel at a time?

Self-attention costs O(N²) in the number of tokens. A 224×224 image is ~50k pixels; attention over those is out of reach. 16×16 patches give only ~196 tokens — computable, and each token carries a local piece of structure into the bargain. It is an engineering compromise between resolution and the cost of attention.

What to read in the original

Read the essentials — the patch embedding and the "you need large data" conclusion; skim the transfer tables.