Era 1 · Origins · 1986

7 Backpropagation

Learning Representations by Back-Propagating Errors · Rumelhart, Hinton & Williams · Nature
🟥 read in full~1–2 h (work it out by hand)original ↗
The gist in 20 seconds. The paper that taught the field to train multilayer networks and made backprop famous. The forward pass computes the output, the backward pass propagates the error by the chain rule, giving a gradient for every weight. And it showed the main thing: hidden layers learn useful internal features on their own. This is the direct answer to Minsky and Papert, and the start of modern DL.

Context

After "Perceptrons" (#4) the field is waiting for a way to train networks with hidden layers — only they get past the linearity limit (XOR). Technically backprop is already known (Linnainmaa/Werbos, #5), but it is this short note in Nature that makes it convincing and sets off the connectionist wave.

The idea and the mechanism

A network of layers of neurons with a smooth nonlinearity (the sigmoid σ) — smoothness is needed so that a derivative exists. We minimize the error, for instance the squared error:

E = 12 Σo (yo − to)2

Training is gradient descent, and the gradient comes from backprop.

Forward pass. For each neuron compute the weighted input netj = Σi wij xi and the activation xj = σ(netj), layer by layer up to the output.

Backward pass. Introduce an "error signal" δj = ∂E / ∂netj. For an output neuron it comes straight from the error; for a hidden one it is assembled from the δ of the neurons in the next layer (that is exactly the chain rule):

δo = (yo − to) · σ′(neto)      δj = σ′(netj) · Σk wjk δk

Once you know δ, the gradient for any weight is just the product of the "error signal from above" and the "activation from below", after which we take a descent step:

∂E / ∂wij = δj · xi      wij ← wij − η · ∂E / ∂wij

The whole "magic" of training a network is the chain rule plus a gradient descent step, arranged into a single backward pass (see #5).

calculus Where the δ rule comes from — a derivation from the chain rule

Define δj ≡ ∂E/∂netj — the sensitivity of the error to the weighted input of neuron j. The error E depends on netj only through the neurons k of the next layer that j feeds. By the chain rule we sum over all such k:

δj = ∂E∂netj = Σk ∂E∂netk · ∂netk∂netj = Σk δk · ∂netk∂netj

Since netk = Σj wjk σ(netj), the derivative of a single term is:

∂netk∂netj = wjk · σ′(netj)

Substitute it back — and out comes the recurrence for the backward pass:

δj = σ′(netj) · Σk wjk δk    ∎

The base of the recursion is the output layer, where E depends on neto directly through yo = σ(neto):

δo = ∂E∂yo · ∂yo∂neto = (yo − to) · σ′(neto)

The gradient for a weight is the last step of the chain rule, where ∂netj/∂wij = xi:

∂E∂wij = δj · xi

The upshot: one backward pass computes δ for every neuron (reusing the δ of the next layer), and from δ you get all the gradients at once. The cost is on the order of a single forward pass, whatever the number of weights. That is what makes training deep networks computationally realistic.

NumPy Implementation: forward + backward for a 2-layer network
import numpy as np
sig = lambda z: 1 / (1 + np.exp(-z))

def forward(x, W1, W2):
    h = sig(x @ W1)             # hidden layer
    y = sig(h @ W2)             # output
    return h, y

def backward(x, h, y, t, W2):
    dy = (y - t) * y * (1 - y)       # δ of the output layer
    dh = (dy @ W2.T) * h * (1 - h)   # δ of the hidden layer (chain rule)
    gW2 = h.T @ dy                   # ∂E/∂W2 = δ · activation
    gW1 = x.T @ dh
    return gW1, gW2                   # step: W -= lr * gW
input x hidden layer (learns features) output y forward pass → ← backward pass: δ
A multilayer network. Values travel forward (gray), the error signal δ travels back (purple). The hidden layer forms intermediate features that nobody specified explicitly.
Analogy. An organization after a failure. The final error (at the output) is handed down the hierarchy by the top manager: each middle manager gets their share of the "blame" (δ) from the units below them and passes it further down. Having received their share, every employee adjusts exactly their own behavior (weight) — in proportion to how much they influenced the outcome.

Why it matters

A direct answer to Minsky and Papert: multilayer networks are trainable and do solve XOR. But the key discovery runs deeper — representation learning: on tasks like symmetry detection the hidden neurons learn meaningful internal features by themselves. This is the core of all of deep learning: the network does not merely classify, it constructs intermediate representations suited to the task.

Backprop becomes the universal training engine — from LeNet (#9) to GPT, everything is trained with it. The limits of that era (vanishing gradients in deep and recurrent networks, not enough data and compute) will surface later and be treated with ReLU, LSTM (#11), ResNet (#27) and GPUs.

Connections

The mathematical core is that same reverse-mode automatic differentiation, discovered by Linnainmaa (1970) and applied to networks by Werbos (1974). The contribution of the 1986 paper is not inventing the algorithm but demonstrating that it can actually train multilayer networks to useful effect, plus a high-profile publication in Nature. To understand backprop is to understand that it is simply an efficient chain rule on the graph of the network.

Minsky and Papert showed the impotence of the single-layer perceptron (it cannot do XOR) and pessimistically assumed the same for multilayer ones. Backprop taught the field to train hidden layers, which build nonlinear features and solve XOR without breaking a sweat — that directly refutes the pessimism and ends the "first AI winter".

→ leads to9. LeNet / CNN

As soon as multilayer networks became trainable, backprop was applied to a convolutional architecture on raw pixels. LeNet is backprop + structural constraints (locality, weight sharing); the same training engine, but now learning a hierarchy of visual features from edges up to digits.

↔ complemented by23. Adam

Backprop answers the question "which way to move the weights" (it gives the gradient), but not "with what step". That is the optimizer's job: from plain SGD to Adam, which adds momentum and per-parameter adaptive steps. The division of labour "gradient (backprop) + step rule (optimizer)" survives to this day in every trainable model.

Questions worth asking

If backprop is just the chain rule, known since the 1970s (#5), why is the 1986 paper considered the turning point?

Because the contribution was not the invention but the demonstration and the publicity. Linnainmaa and Werbos gave the algorithm itself, but this work showed empirically that you can use it to train multilayer networks to good effect — and that hidden layers learn meaningful internal features (representation learning). Plus a publication in Nature and a field that had ripened: interest, growing compute.

"Who discovered it" and "who made the idea work and made it known" are often different people in science. Backprop is a textbook case.

The loss function is non-convex — won't descent get stuck in a bad local minimum?

In practice almost never, and that puzzled people for a long time. In high-dimensional spaces the overwhelming majority of critical points are saddles, not bad minima, and SGD escapes them (the noise helps). Heavily overparameterised networks have whole connected manifolds of good solutions, and empirically different minima give almost the same loss.

The global optimum is not guaranteed, but a "good enough" one is found reliably — one of the empirical puzzles that make deep learning work.

Why a sigmoid rather than a hard threshold, as in the perceptron (#3)?

Backprop needs a derivative, and the derivative of a step function is either zero or undefined — the gradient does not flow, there is nothing to train with. A smooth sigmoid fixes that.

Later it turned out that the sigmoid itself saturates at the ends (the vanishing gradient in deep networks), and it was pushed aside by ReLU — piecewise linear, with derivative 0 or 1, which does not saturate and sharply sped up training of deep networks.

Is that how the brain learns? Is backprop biologically plausible?

Probably not — at least not literally. The backward pass requires the same synaptic weights to be used "in reverse" (the weight transport problem), and no single global error signal is visible in the brain.

There are more plausible alternatives (feedback alignment, predictive coding) that approximate backprop with local rules. But as an engineering tool backprop is under no obligation to copy biology — it simply works.

What to read in the original

The note is short (4 pages) — worth reading in full, and work through the derivation of δ for one hidden layer by hand. That is the most effective way to understand once and for all that training a network is not magic but the chain rule plus gradient descent.