Era 2 · Foundations · 2006

14 Deep Belief Nets

A Fast Learning Algorithm for Deep Belief Nets · Hinton, Osindero & Teh · Neural Computation
🟦 this write-up is enough~1 horiginal ↗
The gist in 20 seconds. A deep network can be trained greedily, layer by layer, without supervision (a stack of RBMs), and then fine-tuned. The method gave deep networks a good initialization and gave them back their respectability — the paper credited with starting "deep learning". The trick itself was later superseded, but its historical role is enormous.

Context

2006: deep networks are formally trainable by backprop, but in practice they train badly — vanishing gradients, poor minima, not enough data; they are considered impractical. Hinton and his co-authors offer a way around it and revive the field's interest.

The idea and the mechanism

A deep generative network built as a stack of RBMs (restricted Boltzmann machines). Train greedily, layer by layer and without supervision: the first RBM models the data; its hidden activations become the "data" for the second RBM; and so on. Then the whole network is carefully fine-tuned (supervised backprop, for instance). Layer-wise pre-training gives the weights a good initialization, from which fine-tuning does converge.

probability RBMs and the contrastive divergence trick

An RBM defines a joint distribution over visible units v and hidden units h through an energy:

P(v, h) ∝ e−E(v,h),   E = −v⊤Wh − a⊤v − b⊤h

We want to maximize the likelihood of the data. The gradient of the log-likelihood has an elegant but intractable form — the difference of two expectations:

∂ log P(v)∂W = ⟨v h⊤⟩data − ⟨v h⊤⟩model

The second expectation requires sampling from the model itself (expensive MCMC). Contrastive Divergence approximates it with a single Gibbs step starting from the data (a "reconstruction") — and that is what made training fast. The RBMs are then stacked greedily: each layer learns the distribution of the previous layer's activations.

NumPy One step of contrastive divergence (CD-1)
import numpy as np
sig = lambda z: 1 / (1 + np.exp(-z))

def cd1(v, W, eta=0.1):                # training a single RBM
    h  = (sig(v @ W) > np.random.rand(W.shape[1])).astype(float)  # data
    v2 = sig(h @ W.T)                  # reconstruction of the visible layer
    h2 = sig(v2 @ W)
    W += eta * (np.outer(v, h) - np.outer(v2, h2))   # data − reconstruction
    return W
data (visible layer) RBM 1 → h₁ RBM 2 → h₂ greedily, layer by layer: 1) train RBM 1 on the data 2) train RBM 2 on the activations h₁ 3) …then fine-tune it all with backprop
The deep network is assembled from the bottom up: each RBM learns the distribution of the previous layer's activations, giving a good initialization for the final fine-tuning.
Analogy. Building a skyscraper not all at once but floor by floor, letting each one set before laying the next. Layer-wise pre-training is "letting the foundation cure": by the time fine-tuning starts, the weights already sit somewhere sensible, and backprop does not fall into a bad minimum.

Why it matters

The paper credited with starting the modern "deep learning" era: it showed that deep networks can be trained and that they give SOTA results (on MNIST), which handed them back their respectability. The Hinton–Bengio–LeCun group held together around it (on a CIFAR grant) through the lean years. A caveat: the method itself (RBMs + layer-wise pre-training) was soon superseded — ReLU, better initialization, BatchNorm and an abundance of data made it possible to train deep networks with plain backprop.

Connections

← builds on6. Hopfield network

The RBM is a relative of the Hopfield network and the Boltzmann machine: the same energy-based models out of statistical physics, but stochastic and trainable. The "energy → probability → learning" line runs straight from here.

→ leads to16. AlexNet

DBNs kindled the belief that deep networks are trainable and pulled a community together. Six years later AlexNet proved it without any RBMs — straight backprop on a GPU. Pre-training turned out to be a crutch that stopped being needed, but it is what carried the field through the dark period.

The DBN answered "how do you get training of a deep network off the ground at all?" with pre-training. BatchNorm (along with good initialization and ReLU) solved the same problem differently — by stabilizing the training process itself, which made layer-wise pre-training unnecessary.

Questions worth asking

If pre-training helped so much, why was it abandoned?

Because the reason it was needed — deep networks training badly — was removed. ReLU (no saturation), sensible initialization (Xavier/He), BatchNorm and large datasets let the gradient flow through depth directly. Once backprop worked head-on, the detour was redundant.

How is RBM pre-training different from autoencoder pre-training?

Both are unsupervised initialization through reconstruction, and the "a layer learns a representation of its input" idea is shared. An RBM is a probabilistic energy-based model trained by contrastive divergence; an autoencoder is a deterministic network minimizing reconstruction error directly with backprop. Autoencoders are simpler and quickly displaced RBMs as the way to pre-train — until they too became unnecessary.

Did this paper "coin" the term deep learning?

No — the term is older. It popularized the modern framing and became the symbolic start of the era, but "deep learning" was in use before that. Crediting it with coining the term is a common mistake; "the paper that brought depth back into the mainstream" is the accurate description.

What to read in the original

This write-up is enough — the historical role matters more than the mechanics of RBMs, which are niche today. If you want to dig, look at the idea of greedy layer-wise training and at contrastive divergence (the math box).