52 Mixtral
Context
MoE (#31) promised scale without a rise in compute, but there were no strong open MoE LLMs. Mistral AI releases Mixtral 8×7B with open weights.
The idea and the mechanism
In every Transformer layer the FFN is replaced by an MoE of 8 experts; a trained router picks 2 experts (top-2) for each token. That is ~47B parameters in total, but only ~13B active per token. It beat Llama-2-70B and GPT-3.5 at less active compute.
optimization Active vs total parameters: where the saving is
The FFN layer is replaced by E=8 experts. The router gives weights, we take the top 2, renormalize and mix:
Do the arithmetic. Total parameters (what has to sit in memory): all 8 experts across all layers ≈ 47B. Active ones (computed per token): only 2 experts ≈ 13B. In other words:
Quality tracks the capacity (47B), while inference speed tracks the active part (13B). This is the same sparse-MoE idea from #31, carried through to a working open LLM: "a lot of knowledge, little arithmetic per token".
PyTorch Mixtral's MoE FFN (top-2 of 8)
import torch
def mixtral_ffn(x, gate, experts): # 8 experts, top-2
w, idx = gate(x).softmax(-1).topk(2, -1)
w = w / w.sum(-1, keepdim=True) # renormalize the weights
out = sum(w[:, j:j+1] * experts[idx[:, j]](x) for j in range(2))
return out # 2 of 8 active → ~13B of 47B parameters per token
Why it matters
It proved to the open community that sparse MoE is practical and made it mainstream among open models; it locked in the engineering pattern of "many total parameters, few active ones", which DeepSeek (#53) then scales further.
Connections
Mixtral is the sparse MoE of 2017 finally grown into a strong open LLM: every layer's FFN is an MoE of 8 experts with top-2 routing. The "Outrageously Large" idea at a practical scale.
DeepSeek-V3 takes the same principle further still — 671B parameters, ~37B active — plus routing improvements (load balancing without an auxiliary loss). A straight line from Mixtral to frontier MoE.
Mixtral is built on a LLaMA-like architecture (RMSNorm, RoPE, SwiGLU) with the dense FFN swapped for an MoE. The open ecosystem LLaMA set up is the soil Mistral grew in.
Questions worth asking
If only 13B is active, why can't you just run it on hardware sized for a 13B model?
Because all 47B parameters have to be in memory — you cannot know in advance which 2 experts the next token will need. Memory is counted in total parameters, not active ones. So MoE saves compute, not VRAM: to run Mixtral you need hardware sized for 47B, and the speed will be that of a ~13B model. It is a different point on the trade-off curve, not a free lunch.
Different tokens go to different experts — doesn't that break coherence within a batch?
It creates an engineering headache: within a batch, tokens spread unevenly across experts, and you have to gather and scatter them efficiently (all-to-all communication) or part of the GPU sits idle. Hence the capacity factor (a cap on tokens per expert) and the elaborate load balancing. MoE is powerful, but its systems implementation is markedly harder than a dense model's.
Do Mixtral's experts specialize by topic (code, language, math)?
The analysis says: barely. How tokens are distributed across experts correlates weakly with topic or domain; the routing picks up syntactic and surface features instead, and is fairly uniform. The appealing intuition of "the medicine expert, the code expert" does not hold up in practice — the router learns whatever lowers the loss, not a human-readable division of labour.
What to read in the original
Read the essentials — the routing (top-2), the difference between active and total parameters, and the effect on memory vs speed.