40 CLIP
Context
Vision classifiers were trained on a fixed set of classes with expensive hand labeling, and they did not transfer to new ones. OpenAI ties vision to language.
The idea and the mechanism
Two encoders — one for the image, one for the text — are trained JOINTLY on 400 million web (image, caption) pairs, contrastively: within a batch, maximize the similarity of the matching pairs and minimize it for all the mismatched ones. What forms is a shared embedding space in which semantically corresponding images and texts sit close together.
probability The contrastive loss (InfoNCE) and zero-shot
In a batch of N pairs the encoders produce image embeddings Ii and text embeddings Tj (L2-normalized). The similarity matrix is Sij = Ii·Tj/τ. A symmetric contrastive loss makes the diagonal (the matching pairs) large and everything off the diagonal small:
where the labels are simply the identity (image i corresponds to text i). Each matching pair is pulled in, the N−1 mismatched ones in the same row and column are pushed away.
Zero-shot. To classify an image, compare its embedding against the embeddings of the texts "a photo of a {class}" and take the nearest one — no fine-tuning on those classes at all. That is how a fixed set of labels gets replaced by any text description.
PyTorch The CLIP contrastive loss
import torch, torch.nn.functional as F
def clip_loss(img_emb, txt_emb, tau=0.07):
I = F.normalize(img_emb, dim=-1)
T = F.normalize(txt_emb, dim=-1)
logits = I @ T.T / tau # (N, N) similarities
labels = torch.arange(len(I)) # diagonal = the matching pairs
return (F.cross_entropy(logits, labels) +
F.cross_entropy(logits.T, labels)) / 2
Why it matters
Natural language as a flexible supervisor in place of fixed labels; CLIP matched ResNet-50 on ImageNet zero-shot. CLIP encoders went on to underpin text-to-image (DALL·E 2, the conditioning in Stable Diffusion) and multimodal systems.
Connections
CLIP's text encoder is a descendant of Transformer language pre-training. The idea of "represent text as a vector that encodes meaning" (BERT/word2vec) is the necessary half that CLIP ties to the visual one.
The same bet on scale of data, one level up: instead of 14M hand-labeled images, 400M web pairs where the labeling is natural language. The data-as-the-engine line evolving from curated annotation to web scale.
CLIP's shared text-image space is the mechanism by which text-to-image understands a request: conditioning in Stable Diffusion leans on CLIP-style text embeddings. Without a language↔vision link, generation "from a description" would be impossible.
Questions worth asking
Why does contrastive training scale better than predicting captions?
Generating an exact caption is orders of magnitude harder than telling the right pair from the wrong ones. The contrastive objective only asks for matching, not for producing text, so it learns more efficiently on noisy web data. CLIP compared both approaches deliberately and chose contrastive for training speed — fewer demands, the same signal.
Zero-shot is sensitive to how the prompt is worded ("a photo of a …") — is that reliable?
Not very. Quality depends noticeably on the template: "a photo of a {class}" works better than the bare word, and an ensemble of many templates works better still. This is the same prompt engineering as in LLMs. CLIP is powerful, but its zero-shot is a brittle interface that needs the wording tuned.
400M web pairs — what about bias and noise?
A serious problem. Web captions are noisy and slanted; CLIP inherits social stereotypes and does poorly on rare or non-European concepts. There are known and disturbing failures when classifying people. So the makeup and provenance of CLIP's data is a point of criticism, and CLIP itself is an example of how scale buys power but also replicates the internet's prejudices.
What to read in the original
Read the essentials — the contrastive setup and the zero-shot prompt; skim the enormous eval sweep.