55 Structured Output
Context
An LLM is a distribution over text, but applications (agents, tool use, API pipelines) need machine-readable output: valid JSON matching a schema exactly. Simply asking it to "reply in JSON" is unreliable: the model drops a comma, adds prose, muddles types. On large schemas prompt-only gets roughly a third of responses valid — unacceptable in production. What you need is a guarantee, not a hope.
The idea and the mechanism
1. Function calling (2023). The model is fine-tuned to emit {"name": …, "arguments": {…}} given a function description and its parameter schema. This made tool use practical — but validity is still probabilistic: a trained model usually hits the schema, but is not obliged to.
2. Constrained / guided decoding (2023–24). The guarantee comes not from training but from the decoding stage. You run an automaton that tracks the already-generated prefix against the grammar; at each step you mask the logits of every token that would make the output invalid, and sample only from what is allowed. The structure is valid by construction. Outlines reduces a regex or grammar to a finite automaton and precomputes an index: state → the set of allowed tokens (otherwise you would have to scan the whole vocabulary at every step). XGrammar takes a context-free grammar, compiles it into a byte-level pushdown automaton and caches the masks. OpenAI Structured Outputs (2024) compiles a JSON schema into a grammar and so guarantees 100% conformance (plus training on schemas).
automata · decoding Masking logits, and why token alignment is the hard part
The core. Let the automaton be in state s, and let \( \mathcal{A}(s) \subseteq V \) be the set of tokens after which a valid completion is still possible. The model gives logits \( z \in \mathbb{R}^{|V|} \); we mask and sample only from the allowed set:
Forbidden tokens get \(-\infty\) → zero probability after the softmax; the distribution is renormalized over the valid ones. The chosen token moves the automaton to a new state \(\delta(s,x_t)\). Which kind of automaton you need depends on the grammar:
A flat format is described by a regular language (a finite DFA). But JSON is recursive — nested, balanced {} and [] do not form a regular language, so you need a pushdown automaton: the stack keeps count of the nesting depth.
The hard part is token alignment. The grammar is defined over characters/bytes \(\Sigma\), while the model generates tokens V (BPE subwords). One token covers several characters, and one string has many tokenizations — there is no direct correspondence. So you cannot simply "allow the right characters": for each state you have to work out the allowed tokens. Hence the precomputed index:
The cost. Checking every token for validity naively is \(O(|V|)\) per step (a vocabulary of ~10⁵). Outlines' index makes it \(\approx O(1)\) (plus a cheap vectorized mask application). XGrammar goes further: it splits tokens into context-independent ones (validity is clear from position alone — precomputable, >99% of them) and context-dependent ones (the whole stack matters), caching masks by the top of the PDA stack. The upshot — <50 µs per token, negligible against 10–50 ms of inference.
Python Constrained decoding: masking logits against an automaton
def constrained_decode(model, automaton, idx): # idx: state -> set(token_id)
s, out = automaton.start, []
while not automaton.is_accept(s):
z = model.logits(out) # logits over the vocabulary V
allowed = idx[s] # precomputed index: tokens allowed from s
z = mask_fill(z, allowed, NEG_INF) # forbidden -> -inf (zero probability)
tok = sample(softmax(z)) # sample ONLY from the valid ones
out.append(tok)
s = automaton.step(s, tok) # automaton transition (regex->DFA, JSON->PDA+stack)
return detokenize(out) # guaranteed valid structure
Why it matters
This is the foundation under all modern tool use, agents and dependable API pipelines: without a guarantee of valid output you cannot build systems where a program parses the LLM's answer. "Structured output", which everyone uses today (OpenAI, Anthropic tool use, vLLM/TGI/llama.cpp with grammars), is exactly this line of work: training on schemas + constrained decoding. The broader lesson: the reliability of LLM systems is often born not in the model itself but in the decoding layer around it.
Connections
Constrained decoding wedges itself in exactly where an autoregressive Transformer turns logits into a token — at the softmax over the vocabulary. Masking logits is possible precisely because generation goes token by token: before each sample you can zero out the forbidden options. Without the step-by-step nature of #32 there would be no such lever.
Function calling is alignment to a tool: the model is fine-tuned to follow call schemas (a relative of the SFT/RLHF in #44). Training makes the output usually valid and sensible; constrained decoding adds the guarantee. Full structured output is the sum of the two: a trained model plus a mask.
CoT wants free reasoning out loud; a strict format suppresses it. The study "Let Me Speak Freely?" (2024) showed that a hard JSON mode hurts reasoning — the model is forced to produce the answer before it has "thought". The practical fix is a free-form CoT field first, format afterwards (the order of fields in the schema decides it).
Questions worth asking
If the mask guarantees valid JSON — why not turn constrained decoding on always?
Because the mask distorts the distribution: zeroing out some tokens and renormalizing can push the model into regions it considers unlikely, and impose structure before it has finished thinking. "Let Me Speak Freely?" (2024) showed a measurable drop in reasoning under a strict format — largely because of the reordering: the answer has to be named before the reasoning. Mitigations: a free reasoning field first, strict fields after (the order in the schema), or generate in natural language and convert to the format in a separate step. A format guarantee is not free.
Why is this harder than "allow only the right characters"?
Because of token alignment. The grammar lives over characters/bytes, the model over BPE tokens: one token glues several characters together, and the same string can be tokenized in different ways. So "allow the character {" does not translate directly into "allow this token" — for every automaton state you have to work out which (multi-character) tokens keep the grammar intact. Hence the precomputed index (Outlines) and the byte-level PDA with its split into context-independent tokens (>99%, precomputable) and context-dependent ones (XGrammar). That is where the real engineering depth of the topic sits.
Regex is enough for many formats — why does JSON specifically need a stack (PDA)?
JSON is recursive: an object can contain objects, arrays nest arbitrarily, brackets must balance. The language of balanced brackets is the classic example of a non-regular one (a finite automaton cannot recognize it: you have to "remember" arbitrary depth). Hence a pushdown automaton with a stack, which counts the open brackets and demands they be closed. For flat formats (a date, an enum, a number) regex/DFA is enough — but the moment the structure nests, you cannot do without a stack.
What to read in the original
Read selectively. Outlines (Willard & Louf, arXiv 2307.09702) — the key idea of the "FSM state → allowed tokens" index and the reduction to an automaton. XGrammar (arXiv 2411.15100) — the byte-level PDA, the split of tokens into context-(in)dependent, the mask cache. The OpenAI Structured Outputs post — how a JSON schema is compiled into a grammar for 100% conformance. And do read "Let Me Speak Freely?" (arXiv 2408.02442) — an honest account of what a strict format costs reasoning.