Era 4 · Architectures and scale · 2016

30 WaveNet

WaveNet: A Generative Model for Raw Audio · van den Oord et al. · DeepMind
🟦 this write-up is enough~30–40 minoriginal ↗
The gist in 20 seconds. It generates raw audio autoregressively — one sample at a time. The key is dilated causal convolutions: the receptive field grows exponentially with depth, covering thousands of samples with no recurrence at all. A new bar for the naturalness of speech synthesis.

Context

Speech synthesis (TTS) sounded robotic (concatenative and parametric methods). DeepMind goes at it radically — generate the audio at the level of individual samples.

The idea and the mechanism

An autoregressive model predicts each next audio sample (at 16 kHz that is 16000 values a second) from all the previous ones. To cover a long context without recurrence it uses dilated causal convolutions: causal (the convolution never looks into the future — correct for generation), dilated (the gaps between the filter's taps grow).

signal processing Why the receptive field grows exponentially

An ordinary convolution with a filter of width k, stacked L layers deep, covers L(k−1)+1 points — linear in depth. A dilated convolution skips d−1 points between the filter's taps. If the dilation doubles layer by layer, d = 1, 2, 4, …, 2L−1, the receptive field becomes:

R = 2L  (grows exponentially with depth)

So about 10 layers cover ~1000 samples, while the computation grows only linearly with depth. Causal means the output at time t depends only on inputs ≤ t (implemented by shifted padding) — a hard requirement for autoregressive generation, or the model would be peeking at a future that does not exist yet at inference time.

PyTorch Dilated causal convolution
import torch.nn as nn
import torch.nn.functional as F

class CausalConv(nn.Module):
    def __init__(self, C, d, k=3):
        super().__init__()
        self.pad = (k - 1) * d                  # shift, so it cannot see the future
        self.conv = nn.Conv1d(C, C, k, dilation=d)
    def forward(self, x):
        return self.conv(F.pad(x, (self.pad, 0)))
# a stack of dilation = 1,2,4,8,… → the field grows exponentially
input (samples) dilation 2 dilation 4 the receptive field doubles with every layer
Each next layer, with a larger dilation, reaches twice as far. A handful of layers cover an enormous context without any recurrence.
Analogy. To take in a whole street you do not have to walk past every house — you just go up the floors: from the first you see the houses next door, from the second the whole block, from the tenth the entire district. Dilation is that climb: with every layer the view doubles, while the number of steps (the computation) barely grows.

Why it matters

It showed that autoregressive generation works for audio too (MOS 4.21 against ~3.7 for the earlier methods), and it shaped the generative models for sound and music that followed. One caveat: the original is horribly slow (one sample at a time); the production version (parallel WaveNet) came later and went into Google Assistant.

Connections

← builds on9. LeNet / CNN

WaveNet is convolution carried over from images to one-dimensional audio and made causal. The same machinery (filters, receptive fields), with the dilation trick added to cover a long stretch of time.

↔ alternative11. LSTM

The same problem — long context in a sequence — with a different answer: an LSTM drags it along recurrently (step by step), WaveNet uses a stack of dilated convolutions (parallel over time during training). A foretaste of how the Transformer would push recurrence out of language too.

WaveNet established autoregressive generation "a piece at a time". Diffusion offers a different route to generation (iterative denoising instead of sample-by-sample prediction), partly curing WaveNet's disease — the agonising slowness of sequential generation.

Questions worth asking

Why is sample-by-sample generation so agonisingly slow, and how was it worked around?

Autoregression is sequential by construction: to emit sample t you need sample t−1 — at 16000 a second that is tens of thousands of forward passes per second of audio. The way round it was distillation into a parallel model (Parallel WaveNet): a fast non-sequential generator was trained to imitate the slow WaveNet. The same "quality vs the speed of autoregression" dilemma resurfaces in LLMs (hence speculative decoding).

Why model audio at the level of samples rather than spectrograms or phonemes?

Earlier TTS worked on intermediate representations and lost naturalness at the joins. WaveNet generates the raw waveform directly, throwing away none of the fine structure (breath, overtones) — hence the jump in naturalness. The price is computational: modeling 16000 points a second costs more than a dozen phonemes.

Are dilated convolutions a WaveNet invention?

No — dilation (atrous convolution) had been used earlier in image segmentation, to widen the receptive field without losing resolution. WaveNet's contribution is joining dilated and causal for autoregressive audio generation, and showing that thousands of samples can be covered that way. "Novelty" in ML is often a well-judged transplant of a trick from a neighbouring field.

What to read in the original

The idea is simple — the write-up is enough. The main things to take away are the dilated causal convolution and the way the receptive field grows exponentially (the math box).