Era 5 · The LLM era · 2019

36 GPT-2

Language Models are Unsupervised Multitask Learners · Radford, Wu, Child, Luan, Amodei, Sutskever · OpenAI
🟧 read selectively~1 horiginal ↗
The gist in 20 seconds. The same decoder-only GPT, but up to 1.5B parameters on quality web text. The central claim: a large enough LM, trained to do nothing but predict the next word, solves tasks ZERO-SHOT — provided you state the task in the prompt. A hint that scale on its own gives rise to general abilities.

Context

GPT-1 was fine-tuned per task. OpenAI checks: what if you simply make the model and the data MUCH bigger?

The idea and the mechanism

The same decoder-only Transformer, but up to 1.5B parameters, trained on WebText (a curated web corpus). No mechanism changes — scale only. And it turns out: the model does translation, question answering and summarization ZERO-SHOT, with no fine-tuning, as long as the task is expressed in the text of the prompt.

probability Why an ordinary LM solves tasks zero-shot

A language model learns one thing — the distribution P(x) over text. But if the task and the input are expressed as text, the conditional probability of the continuation is the answer:

P(output | "task: description; input: x ⟹")

Given "Translate to French: cat ⟹", for instance, the model continues with "chat". Where does that skill come from? A web corpus contains masses of implicit demonstrations of tasks — translations, questions and answers, summaries, code with comments. In minimizing the language-modeling loss on such text, the model incidentally learns to perform those tasks. Hence the title: "unsupervised multitask learners" — multitasking arises by itself, as a by-product of next-word prediction on varied data.

Python Zero-shot through the prompt
# the task is stated in the prompt itself, no fine-tuning and no examples
prompt = "Translate to French: cat =>"
out = lm.generate(prompt)        # the model continues: " chat"

prompt = "Summary: <long text> TL;DR:"
out = lm.generate(prompt)        # the model continues with a summary
one LM "translate: …" "question: …" "TL;DR: …" answer
One and the same model solves different tasks — all that differs is how the task is worded in the prompt. There is no fine-tuning.
Analogy. A well-read generalist who was never coached for any particular task. Ask them "how do you say cat in French?" and they answer, because they have come across translations in their reading. Ask them for a summary and they manage, because they have seen thousands of abstracts. They never studied these tasks separately; the ability settled in on its own, out of plenty of reading.

Why it matters

A shift towards the idea of a general-purpose LM steered by a prompt — which GPT-3 would confirm outright, and which would become ChatGPT. Famous for its staged release over fears of misuse ("too dangerous to release" → later "no evidence of misuse").

Connections

← scales up35. GPT-1

Same architecture, same recipe — only more model and more data. GPT-2 is the experiment "what does pure scale buy you", and the answer turned out to be a qualitative jump (zero-shot), not merely a quantitative one.

→ leads to38. GPT-3

GPT-2 hinted at zero-shot; GPT-3 pushed the scale two orders of magnitude further and opened up few-shot in-context learning. A straight ladder of scale: 1.5B → 175B, and with it the move from "a hint of abilities" to "a general few-shot solver".

← justified by37. Scaling Laws

GPT-2's observation that "bigger = qualitatively better" was soon formalized as scaling laws: the loss falls as a predictable power law in scale. That turned "let's make it bigger" from an intuition into an engineering strategy.

Questions worth asking

"Too dangerous to release" — a real threat, or marketing?

Contested. OpenAI justified the staged release by the risk of mass-produced disinformation; critics called it inflated and a PR move. In the event no catastrophe happened, and the full model was published about nine months later. The episode matters as the first loud precedent in the debate about "responsible release" of powerful models — a topic that has only sharpened since.

Zero-shot works, but badly — why is it still counted as a breakthrough?

Because what matters is not the quality level but the fact: a model that nobody taught to translate translates — purely out of language modeling. That changed the question from "how do we train it for the task" to "how does scale give rise to abilities". The weak zero-shot of 2019 is the forerunner of GPT-3's strong few-shot and of instruct models; the breakthrough is in the direction, not in the metric.

Corpus quality (WebText) — how critical is it?

Very. WebText was assembled from links posted on Reddit with positive karma — a crude but effective quality filter. Junk web text would have made the model worse; curating the data turned out to matter as much as scale. The lesson that "data decides" has only grown stronger since — modern frontier models spend enormous effort on filtering and corpus composition.

What to read in the original

Read the essentials — the zero-shot framing and the role of WebText; the benchmark tables can be skimmed.