43 Chain-of-Thought
Context
LLMs were bad at multi-step problems (arithmetic, logic) — they produced an answer "straight away" and got the intermediate steps wrong. Wei et al. change that with a single trick.
The idea and the mechanism
In the few-shot examples, show the model not just the answer but the INTERMEDIATE STEPS of the reasoning ("let's think step by step"). Then, on a new problem, the model also generates a chain of reasoning BEFORE the answer — and accuracy jumps. Eight CoT examples took PaLM 540B to SOTA on GSM8K, beating even a fine-tuned GPT-3 with a verifier.
probability Why intermediate steps help
A direct answer is P(a | q): all the computation has to fit into a single forward pass, and on a multi-step problem the accuracy is low. CoT introduces an explicit chain of reasoning r:
Generating r has two effects. (1) More computation: every generated token is another forward pass, so the reasoning unfolds over many sequential steps instead of one. (2) Conditioning on intermediate results: the answer rests on sub-conclusions written out explicitly rather than held "in the head". The key observation: the effect is emergent — for small models CoT barely helps (or hurts), while for large ones it gives a big win. The ability to "reason step by step" only appears with scale.
Python A CoT prompt (few-shot with reasoning)
prompt = """Q: Roger has 5 balls. He buys 2 more cans of 3 balls each. How many now?
A: He started with 5. He bought 2 × 3 = 6. So 5 + 6 = 11. Answer: 11.
Q: The canteen had 23 apples, 20 went to lunch, then 6 more were bought. How many apples?
A:"""
out = lm.generate(prompt) # the model writes out the steps, then the answer: 9
Why it matters
In practice — a powerful free trick (CoT → self-consistency, ReAct, scratchpad). Conceptually — the push that started the reasoning line, culminating in models with long "thinking" (o1, DeepSeek-R1, #53) trained to do CoT via RL. One prompt format unlocked a capability the model already had.
Connections
CoT is "advanced few-shot": the same in-context learning interface as GPT-3, only with reasoning steps added to the examples. Without the "learn from examples in the prompt" paradigm there would be nothing for CoT to stand on.
CoT showed that reasoning step by step lifts quality sharply. R1 takes the next step: it trains the model to reason via RL on verifiable rewards, and long chains of thought emerge on their own. Prompting in 2022 grew into a trained ability to reason.
Both are about "spend more compute at solve time". AlphaGo unfolds a tree search; CoT unfolds reasoning in tokens. Two forms of inference-time compute — think longer to solve better.
Questions worth asking
Does the chain of reasoning reflect the model's actual "thinking"?
Not necessarily. It is plausible text that helps arrive at an answer, but it is not guaranteed to describe the internal computation. Sometimes the model produces the right answer off a faulty chain, or "rationalises" a decision it had already made. CoT is a useful tool and a partial window into the reasoning, but not a reliable log of the internal work (the "faithfulness of reasoning" topic).
Why is CoT emergent — why does it not work on small models?
A small model lacks the basic sub-skills (reliable arithmetic, holding context) for a chain of steps to be correct — one bad step wrecks the whole derivation, and CoT actually hurts. A large model is reliable enough at each step, and only then does breaking the problem into steps start to pay. This is an example of "a capability appearing in a jump with scale", which a smooth loss curve predicts poorly (#37).
Can you get the CoT gain without examples in the prompt?
Yes — zero-shot CoT: just append "let's think step by step" and the model unfolds the reasoning by itself (Kojima 2022). And self-consistency amplifies the effect: sample several chains and take the most frequent answer. That is how one trick grew into a whole family of "think longer at inference" methods.
What to read in the original
It is short — read it in full; pay attention to the emergence and to how a simple change of interface unlocks capabilities the model already has.