30 WaveNet
Context
Speech synthesis (TTS) sounded robotic (concatenative and parametric methods). DeepMind goes at it radically — generate the audio at the level of individual samples.
The idea and the mechanism
An autoregressive model predicts each next audio sample (at 16 kHz that is 16000 values a second) from all the previous ones. To cover a long context without recurrence it uses dilated causal convolutions: causal (the convolution never looks into the future — correct for generation), dilated (the gaps between the filter's taps grow).
signal processing Why the receptive field grows exponentially
An ordinary convolution with a filter of width k, stacked L layers deep, covers L(k−1)+1 points — linear in depth. A dilated convolution skips d−1 points between the filter's taps. If the dilation doubles layer by layer, d = 1, 2, 4, …, 2L−1, the receptive field becomes:
So about 10 layers cover ~1000 samples, while the computation grows only linearly with depth. Causal means the output at time t depends only on inputs ≤ t (implemented by shifted padding) — a hard requirement for autoregressive generation, or the model would be peeking at a future that does not exist yet at inference time.
PyTorch Dilated causal convolution
import torch.nn as nn
import torch.nn.functional as F
class CausalConv(nn.Module):
def __init__(self, C, d, k=3):
super().__init__()
self.pad = (k - 1) * d # shift, so it cannot see the future
self.conv = nn.Conv1d(C, C, k, dilation=d)
def forward(self, x):
return self.conv(F.pad(x, (self.pad, 0)))
# a stack of dilation = 1,2,4,8,… → the field grows exponentially
Why it matters
It showed that autoregressive generation works for audio too (MOS 4.21 against ~3.7 for the earlier methods), and it shaped the generative models for sound and music that followed. One caveat: the original is horribly slow (one sample at a time); the production version (parallel WaveNet) came later and went into Google Assistant.
Connections
WaveNet is convolution carried over from images to one-dimensional audio and made causal. The same machinery (filters, receptive fields), with the dilation trick added to cover a long stretch of time.
The same problem — long context in a sequence — with a different answer: an LSTM drags it along recurrently (step by step), WaveNet uses a stack of dilated convolutions (parallel over time during training). A foretaste of how the Transformer would push recurrence out of language too.
WaveNet established autoregressive generation "a piece at a time". Diffusion offers a different route to generation (iterative denoising instead of sample-by-sample prediction), partly curing WaveNet's disease — the agonising slowness of sequential generation.
Questions worth asking
Why is sample-by-sample generation so agonisingly slow, and how was it worked around?
Autoregression is sequential by construction: to emit sample t you need sample t−1 — at 16000 a second that is tens of thousands of forward passes per second of audio. The way round it was distillation into a parallel model (Parallel WaveNet): a fast non-sequential generator was trained to imitate the slow WaveNet. The same "quality vs the speed of autoregression" dilemma resurfaces in LLMs (hence speculative decoding).
Why model audio at the level of samples rather than spectrograms or phonemes?
Earlier TTS worked on intermediate representations and lost naturalness at the joins. WaveNet generates the raw waveform directly, throwing away none of the fine structure (breath, overtones) — hence the jump in naturalness. The price is computational: modeling 16000 points a second costs more than a dozen phonemes.
Are dilated convolutions a WaveNet invention?
No — dilation (atrous convolution) had been used earlier in image segmentation, to widen the receptive field without losing resolution. WaveNet's contribution is joining dilated and causal for autoregressive audio generation, and showing that thousands of samples can be covered that way. "Novelty" in ML is often a well-judged transplant of a trick from a neighbouring field.
What to read in the original
The idea is simple — the write-up is enough. The main things to take away are the dilated causal convolution and the way the receptive field grows exponentially (the math box).