46 DDPM (Diffusion)
Context
GANs (#20) gave sharp images but were temperamental (instability, mode collapse). Ho, Jain and Abbeel make diffusion practical and good.
The idea and the mechanism
Forward (fixed, nothing learned): over T steps, gradually add Gaussian noise until only pure noise is left. Reverse (learned): the network learns to UNDO that step by step. Generation: start from pure noise and denoise iteratively into an image.
probability From the variational bound to a plain MSE on the noise
Forward. Adding noise has a closed form for any step (a property of Gaussians): you can get xt from x0 in one go:
where ᾱt = ∏s≤t(1−βs). Reverse. We learn pθ(xt−1 | xt). The full variational lower bound (ELBO) is unwieldy, but the authors show that it reduces to a surprisingly simple form — predict the added noise:
That is, the network εθ looks at the noised xt and the step t and predicts which noise was mixed in — training is an ordinary MSE. The beauty of it is that a hard generative problem has turned into "guess the noise".
PyTorch A diffusion training step
import torch
def diffusion_loss(x0, model, abar, T):
t = torch.randint(0, T, (len(x0),))
eps = torch.randn_like(x0)
a = abar[t].view(-1, 1, 1, 1)
xt = a.sqrt() * x0 + (1 - a).sqrt() * eps # add the noise in one step
return ((eps - model(xt, t)) ** 2).mean() # predict the noise (MSE)
Why it matters
It gave GAN-quality generation WITHOUT the instability and with full mode coverage (diversity) → diffusion became the dominant paradigm for image generation (and audio, video, molecules). The direct foundation of Stable Diffusion (#47), DALL·E 2, Imagen.
Connections
The same task (image generation), the opposite approach: instead of two networks competing, one network learns to denoise. Diffusion cured the GAN's signature illnesses (instability, mode collapse) and pushed it out of image generation by the early 2020s.
DDPM works in pixels — expensive. Stable Diffusion moves the same math into a compact latent space, putting text-to-image within reach of consumer GPUs. A direct practical continuation.
Both generate iteratively, but differently: WaveNet is autoregressive, one sample at a time; diffusion works on the whole image in parallel, but over many denoising steps. Diffusion partly cures the slowness of sequential generation (though it still needs dozens of steps of its own).
Questions worth asking
Why predict the noise rather than the clean image directly?
Mathematically it is equivalent (given the noise and xt you can recover x0), but numerically ε-prediction gives a better-conditioned problem: the target has unit variance at every step, so the gradients are steadier. Empirically, predicting the noise trains noticeably better than predicting the image — a lucky choice of parameterization.
Diffusion needs dozens or hundreds of denoising steps — does that not kill the speed?
It does, and that is the main price paid against a GAN (a single pass). Hence the accelerators: DDIM (deterministic sampling in fewer steps), distillation into few-step or one-step models, consistency models. So "many steps" is a real weakness, actively being worked on; the quality and robustness of diffusion outweighed the slowness.
Where did the idea come from — was it invented from scratch in 2020?
No: the theoretical basis is non-equilibrium thermodynamics (Sohl-Dickstein, 2015), and the roots run back into score matching and stochastic differential equations. DDPM's contribution was making the approach practical: simplifying it down to ε-prediction and an MSE loss turned an elegant but unwieldy theory into a working recipe at SOTA quality.
What to read in the original
Read it in full — the math (the forward process as a marginal, the simplification to ε-prediction) is both important and lovely; it is the best way to see why "train a denoiser" = "train a generative model".