Where all this is heading
The threads that run through the canon
An additive path rescues the gradient. The same "+1 in the derivative" surfaces three times: the constant error carousel in the LSTM (ct = ct−1 + …, #11), the identity shortcut in ResNet (y = x + F(x), #27) and the residual wrappers in the Transformer (#32). Depth and length were tamed not by a clever activation but by a direct path for the gradient.
Scale — and its limits. From "bigger = better" (#37 Kaplan) to "grow the data alongside the parameters" (#42 Chinchilla), and on into the wall of finite data, which opened up a new axis: spending compute at inference (#58 o1 / test-time). Every time one scaling axis hits its ceiling, the next one turns up.
Alignment keeps getting simpler. RLHF on PPO (#44, #33) → DPO drops the RL (#51) → Constitutional AI drops the human from the labeling (#45) → GRPO drops the critic (#59). The complicated three-stage pipeline is being taken apart piece by piece into cheaper, steadier parts.
The KV cache is the inference bottleneck. A whole branch of Era 6 fights over the same thing: FlashAttention makes computing attention cheaper (#48), vLLM manages the cache (#49), GQA cuts the number of KV heads (#56), MLA compresses each one (#53), prefix caching and CacheBlend reuse the cache across requests (#54). Progress here is classic systems engineering, not new architectures.
An idea waits for its substrate. LeNet's convolutions in '98 (#9) waited for GPUs and data until AlexNet in '12 (#16); the sparse MoE of 2017 (#31) ripened into Mixtral and DeepSeek (#52, #53). The right idea often arrives a decade ahead of the hardware that lets it fire.
Reasoning as a resource. Chain-of-Thought showed that reasoning step by step helps (#43); o1 turned its length into a quality dial and taught it by RL (#58); R1 and GRPO made the recipe open (#53, #59). "Think longer" went from a prompting trick to a learnable, scalable ability.
Reliability lives in the decoding layer. Structured output (#55) is a reminder: part of what makes LLM systems fit for production is born not in the model's weights but in the wrapper around it — the logit mask, the automaton, the verifier.
Where things seem to be heading
Extrapolating from the canon is a thankless business (half the papers here caught their own contemporaries by surprise). But a few directions are visible straight from its last pages:
Test-time compute and reasoning — the newest axis (#58, #59): probably a further shift of "intelligence" out of training and into inference, search and verification, and a fight to make that thinking cheap and dependable.
Agents and tool use — a direct continuation of structured output (#55) and function calling: the LLM as a component of systems that call tools and act over many steps; the bottleneck here is reliability and evaluation, not "intelligence" per se.
Inference efficiency — the KV-cache branch (#48–#57) is not closed: long context, cheap serving and low latency remain an arena for systems engineering.
The data shortage (#42) pushes towards synthetic data, training on the model's own reasoning (as in R1), and multimodality as a new source of signal.
Which of these turns out to be the next "paper #60" — time will tell. A canon is a canon precisely because it keeps being written.