58 o1 / Test-Time Compute
Context
Scaling laws (#37) and Chinchilla (#42) scaled training: more parameters and more tokens. But good data is finite (#42), and leaning on a single lever eventually runs out. Another lever was left almost untouched — spending compute not on training but on inference: letting the model "think" longer about its answer.
The idea and the mechanism
o1 is trained with RL to produce a long internal chain of reasoning — not as a prompting trick (CoT #43) but as a learned ability: writing out steps, checking itself, going back over mistakes, trying different routes. The key observation: quality grows monotonically both with train-time RL and with test-time "thinking" — the more reasoning tokens at inference, the more accurate the answer on hard problems. Snell et al. formalize how to spend the test-time budget optimally (search against verifier models vs adaptive revision of the answer) and show that allocating the budget by prompt difficulty (compute-optimal) is up to ×4 more efficient than naive best-of-N — and that sometimes a small model with test-time compute beats a far bigger one.
optimization · scaling Two budget axes and verifier search
The simplest way to spend test-time compute is to sample N answers and pick the best one by a verifier model V:
There is also a "vertical" way — making one chain longer and better (adaptive revision). Quality is a function of two budgets; for a fixed total, the optimal split shifts towards inference on hard problems:
Snell et al.: naive best-of-N spends the same budget on easy and hard prompts alike; a compute-optimal strategy allocates it adaptively by difficulty — hence the >4× efficiency win. RL (o1), meanwhile, teaches the model to spend its "thinking" usefully rather than merely longer.
Python Test-time scaling: best-of-N with a verifier
def answer(model, verifier, x, N):
ys = [model.sample(x) for _ in range(N)] # N chains of reasoning — this is the test-time compute we spend
return max(ys, key=verifier.score) # pick the best one by the verifier
# compute-optimal: N depends on the difficulty of x (small for easy, large for hard)
# o1 goes further: RL teaches the model to produce ONE long self-checking chain
Why it matters
It shifted the frontier's paradigm from "more parameters" to "think more" and opened up a class of reasoning models (o1, then o3, DeepSeek-R1 and a wave of others). It is a direct answer to running out of data (#42): as the training axis gets harder to scale, a second one appears — inference. For system design it changes the economics: part of the "intelligence" moves out of the model weights and into the inference budget.
Connections
Chinchilla optimized the training compute (parameters vs tokens). Test-time compute adds an orthogonal axis: spending at inference. When training data is scarce, it is the inference axis that gives the next jump.
CoT (2022) showed that reasoning step by step helps — but through the prompt. o1 turns the length and quality of that reasoning into a resource you can scale accuracy with, and teaches it by RL rather than by prompting.
R1 is the public counterpart of the o1 idea: reasoning grown by RL on verifiable rewards. The same line of "test-time reasoning as an ability", but with open weights and a written-down recipe (see #59 GRPO).
Questions worth asking
How is o1's "thinking longer" different from just a long CoT prompt?
CoT (#43) is a prompting trick: we ask the model to reason step by step. o1 is trained (by RL) to generate long chains that are actually useful — to check itself, drop dead-end branches, come back to mistakes. Thinking becomes a learnable ability, and its length becomes a quality dial. The difference is like the one between asking someone to think out loud and training them to think effectively.
Why can a small model plus test-time compute beat a big one?
On hard problems, search, verification and revision at inference add "effective capacity" more cheaply than inflating the parameter count. Snell et al. showed that allocating the test-time budget by difficulty (compute-optimal) is >4× more efficient than best-of-N — and a small model given time to "think" beats a big one that answers straight away. Compute moves from training into inference.
Is there a limit to test-time scaling?
Yes. On easy problems the extra thinking does not help (and sometimes "overthinking" / overcomplicating actively hurts). Inference costs rise — answers become expensive and slow. The gain depends on the quality of the verifier or the reward. And it does not replace training, it complements it. Test-time compute is a powerful lever, but not a free or a universal one.
What to read in the original
Read both in full. OpenAI's "Learning to Reason with LLMs" — the curves of quality growing with both train-time RL and test-time thinking (the central claim). Snell et al. (arXiv 2408.03314) — the formalization of compute-optimal test-time: search against a verifier vs revision, allocating the budget by difficulty, and the comparison with best-of-N.