Era 6 · Generative models and systems · 2023

50 LLaMA

LLaMA: Open and Efficient Foundation Language Models · Touvron, Lavril, Izacard et al. · Meta AI
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. A family of efficient models (7B–65B) trained on PUBLIC data only: the 13B beats GPT-3 175B. Its architectural details (RMSNorm, SwiGLU, RoPE) became the standard for open models. The leaked weights seeded the entire open-source LLM ecosystem.

Context

The strong LLMs were closed (GPT-3/3.5). Meta releases a family of efficient foundation models trained on public data.

The idea and the mechanism

Following the lessons of Chinchilla (#42), they take smaller models (7B–65B) but train them on MORE tokens → a better deal in terms of INFERENCE compute. The 13B LLaMA beats GPT-3 175B on many benchmarks. The architectural details became the standard: RMSNorm (normalization), SwiGLU (the FFN activation), RoPE (rotary positional embeddings).

linear algebra RoPE: position as a rotation of vectors

How do you encode position? RoPE rotates pairs of coordinates in Q and K by an angle proportional to the position. For a pair (x1, x2) at position m:

\[ \begin{pmatrix} x'_1 \\ x'_2 \end{pmatrix} = \begin{pmatrix} \cos m\theta & -\sin m\theta \\ \sin m\theta & \cos m\theta \end{pmatrix} \begin{pmatrix} x_1 \\ x_2 \end{pmatrix} \]

The magic is that the dot product of the rotated qm and kn depends only on the relative position (m−n), not on the absolute ones. That gives you relative positioning "for free" (no separate embeddings) and extrapolates better to lengths not seen in training. Add RMSNorm (normalizing by the root mean square — cheaper than LayerNorm) and SwiGLU (a gated activation, empirically better than ReLU) — these three choices became the de facto standard architecture for open LLMs.

PyTorch RoPE — rotating Q/K by position
import torch
# rotate pairs of dimensions by an angle ∝ position
def rope(x, pos, theta):
    ang = pos * theta                      # the angle grows with position
    x1, x2 = x[..., 0::2], x[..., 1::2]
    return torch.cat([x1*ang.cos() - x2*ang.sin(),
                      x1*ang.sin() + x2*ang.cos()], dim=-1)
# after RoPE: q_m · k_n depends only on (m − n)
LLaMA 13Bmany tokens GPT-3 175Bbigger, but few tokens 13B ≳ 175B cheaper inference
A smaller model trained for longer not only catches up with the giant, it is also cheaper to run — which matters more in practice.
Analogy. RoPE is like the hand of a clock: position in time is given by an angle of rotation. To work out "how much later one event is than another" you look at the difference of the angles, not at absolute time. The model does the same: having rotated token representations by an angle set by their position, it reads the relative distance between words straight off the geometry — naturally, and in a way that carries over to long texts.

Why it matters

The weights (under a research licence) leaked in March 2023 → an open-source explosion: Alpaca, Vicuna and thousands of fine-tunes; it started the wave of open LLMs that Mistral and the rest grew out of. LLaMA-2 and 3 are officially open.

Connections

← applies42. Chinchilla

LLaMA is Chinchilla's lesson made flesh: smaller models on far more tokens. And it goes further — it trains past the Chinchilla point for the sake of cheap inference, which pays for itself over millions of requests.

→ fine-tuned via41. LoRA

Open LLaMA weights + LoRA = an explosion of customization: the community fine-tunes the base with lightweight adapters on consumer hardware. An open model and cheap fine-tuning together are what made iteration this fast.

An open model has to be run somewhere efficiently — vLLM became the standard serving engine for the LLaMA family. Open weights + efficient inference = practical self-hosting.

Questions worth asking

Why was the "weights leak" a turning point rather than just an incident?

Because for the first time the community had a strong base model it could freely fine-tune and run. Before that, open models lagged noticeably behind. LLaMA gave a foundation on which Alpaca, Vicuna and hundreds of projects grew within weeks — which shifted the center of innovation partly out of the labs and into open source. Meta later drew its conclusions and released LLaMA-2 and 3 openly by design.

Why train past the Chinchilla point if it is "optimal"?

Chinchilla optimizes training under a fixed compute budget. But for a deployed model what matters more is the cost of inference (millions of requests). A small model is cheaper to serve, so it pays to "over-train" it with extra tokens: you spend more on training once and save forever on requests. LLaMA deliberately goes against compute-optimal in favor of deploy-optimal.

RMSNorm, SwiGLU, RoPE — are these fundamental improvements or minor hacks?

More like carefully selected empirical improvements: each gives a small gain, but together they are noticeable and cost almost no extra compute. RMSNorm is cheaper, SwiGLU is slightly better than ReLU, RoPE is better for long contexts. Their value is that they became a standard set: almost every modern open LLM uses exactly this trio.

What to read in the original

Read the essentials — the data and architecture choices (RMSNorm, SwiGLU, RoPE) and the "smaller, but trained longer" results.