Era 5 · The LLM era · 2021

40 CLIP

Learning Transferable Visual Models From Natural Language Supervision · Radford, Kim et al. · OpenAI · ICML
🟧 read selectively~1.5–2 horiginal ↗
The gist in 20 seconds. Train two encoders (image and text) on 400M web pairs contrastively: pull the matching pairs together, push the mismatched ones apart. You get a shared embedding space → zero-shot classification from text prompts. Natural language as a flexible supervisor.

Context

Vision classifiers were trained on a fixed set of classes with expensive hand labeling, and they did not transfer to new ones. OpenAI ties vision to language.

The idea and the mechanism

Two encoders — one for the image, one for the text — are trained JOINTLY on 400 million web (image, caption) pairs, contrastively: within a batch, maximize the similarity of the matching pairs and minimize it for all the mismatched ones. What forms is a shared embedding space in which semantically corresponding images and texts sit close together.

probability The contrastive loss (InfoNCE) and zero-shot

In a batch of N pairs the encoders produce image embeddings Ii and text embeddings Tj (L2-normalized). The similarity matrix is Sij = Ii·Tj/τ. A symmetric contrastive loss makes the diagonal (the matching pairs) large and everything off the diagonal small:

L = ½ [CE(softmaxrow(S), I) + CE(softmaxcol(S), I)]

where the labels are simply the identity (image i corresponds to text i). Each matching pair is pulled in, the N−1 mismatched ones in the same row and column are pushed away.

Zero-shot. To classify an image, compare its embedding against the embeddings of the texts "a photo of a {class}" and take the nearest one — no fine-tuning on those classes at all. That is how a fixed set of labels gets replaced by any text description.

PyTorch The CLIP contrastive loss
import torch, torch.nn.functional as F

def clip_loss(img_emb, txt_emb, tau=0.07):
    I = F.normalize(img_emb, dim=-1)
    T = F.normalize(txt_emb, dim=-1)
    logits = I @ T.T / tau              # (N, N) similarities
    labels = torch.arange(len(I))       # diagonal = the matching pairs
    return (F.cross_entropy(logits, labels) +
            F.cross_entropy(logits.T, labels)) / 2
image text encoder encoder shared matching pair→ close (green)
An image and its caption are mapped into one space and pulled together; pairs that do not belong are pushed apart. That is how language becomes a supervisor for vision.
Analogy. Teaching a child not with flashcards carrying fixed labels ("this is #347 in the catalogue") but by showing pictures and saying the captions out loud in ordinary language. Then they can recognize things the catalogue never contained: hearing a new description, they will find the matching picture. CLIP is likewise not tied to a list of classes — it understands any text.

Why it matters

Natural language as a flexible supervisor in place of fixed labels; CLIP matched ResNet-50 on ImageNet zero-shot. CLIP encoders went on to underpin text-to-image (DALL·E 2, the conditioning in Stable Diffusion) and multimodal systems.

Connections

← builds on34. BERT

CLIP's text encoder is a descendant of Transformer language pre-training. The idea of "represent text as a vector that encodes meaning" (BERT/word2vec) is the necessary half that CLIP ties to the visual one.

← inherits from15. ImageNet

The same bet on scale of data, one level up: instead of 14M hand-labeled images, 400M web pairs where the labeling is natural language. The data-as-the-engine line evolving from curated annotation to web scale.

CLIP's shared text-image space is the mechanism by which text-to-image understands a request: conditioning in Stable Diffusion leans on CLIP-style text embeddings. Without a language↔vision link, generation "from a description" would be impossible.

Questions worth asking

Why does contrastive training scale better than predicting captions?

Generating an exact caption is orders of magnitude harder than telling the right pair from the wrong ones. The contrastive objective only asks for matching, not for producing text, so it learns more efficiently on noisy web data. CLIP compared both approaches deliberately and chose contrastive for training speed — fewer demands, the same signal.

Zero-shot is sensitive to how the prompt is worded ("a photo of a …") — is that reliable?

Not very. Quality depends noticeably on the template: "a photo of a {class}" works better than the bare word, and an ensemble of many templates works better still. This is the same prompt engineering as in LLMs. CLIP is powerful, but its zero-shot is a brittle interface that needs the wording tuned.

400M web pairs — what about bias and noise?

A serious problem. Web captions are noisy and slanted; CLIP inherits social stereotypes and does poorly on rare or non-European concepts. There are known and disturbing failures when classifying people. So the makeup and provenance of CLIP's data is a point of criticism, and CLIP itself is an example of how scale buys power but also replicates the internet's prejudices.

What to read in the original

Read the essentials — the contrastive setup and the zero-shot prompt; skim the enormous eval sweep.