2 Hebbian learning
Context
The McCulloch–Pitts neuron (1943) could compute but could not learn — its weights are fixed. The psychologist Donald Hebb goes looking for a biological mechanism that explains how experience changes the brain. His postulate will become the cornerstone of learning theory in neural networks.
The idea and the mechanism
Hebb's postulate: if neuron A systematically takes part in firing neuron B, the connection A→B strengthens. In its simplest form the weight change is proportional to the product of the activities:
The rule is local (it needs only the activities of the two connected neurons, no global error signal) and works unsupervised — the network soaks up the correlation statistics of its inputs. Hebb also introduced cell assemblies: groups of co-firing neurons that get "stitched together" into a stable pattern carrying a concept.
linear algebra Why pure Hebb is unstable — and how it connects to PCA
Take a linear neuron y = w·x and feed its output into the Hebb rule Δw = η y x. Averaging over the data:
C is the covariance matrix of the inputs. This is power iteration: w grows along the leading eigenvector of C — the direction of maximum variance in the data. But the norm of w blows up without bound while it does so.
The cure is Oja's rule (1982), which adds a normalizing term:
Now ‖w‖ → 1, and the neuron converges to the unit leading eigenvector — that is, it computes the first principal component (PCA). The conclusion: Hebb + normalization = learning a representation that picks out the main direction in the data.
NumPy Implementation: Oja's rule → the principal component
import numpy as np
def oja(X, eta=0.01, epochs=50):
w = np.random.randn(X.shape[1]) * 0.01
for _ in range(epochs):
for x in X:
y = w @ x # linear neuron
w += eta * y * (x - y * w) # Hebb + normalization (Oja)
return w / np.linalg.norm(w) # ≈ 1st principal component of X
Why it matters
This is the first rule by which a network changes itself from experience — the move from "a neuron as a gate" to "a neuron as a learning element". Out of it grow Hopfield's associative memory (its weights are set by the Hebb rule), self-organizing maps, and — on the biology side — long-term potentiation (LTP), the neural substrate of memory. The slogan "cells that fire together, wire together" is a late paraphrase (~1992), but it captures the point exactly.
Connections
The MP neuron gives you a unit of computation, but with fixed weights. Hebb adds a mechanism of plasticity to it — the connections change on their own. Together they are the first picture of a "learning network": computation (MP) plus learning (Hebb).
Hopfield takes the Hebb rule literally: the network's weights are set as a sum of outer products of the patterns to be stored. That turns an abstract postulate into a working model of associative memory with provable dynamics.
Two different spirits of learning. Hebb is unsupervised: it changes weights from the correlation of activities, without knowing the "right answer". The perceptron is supervised: it moves weights along an error signal relative to a label. That fork (unsupervised vs supervised) runs through the whole history of ML.
Questions worth asking
If the rule only strengthens connections, why doesn't the network "saturate" — every weight at maximum?
Pure Hebb really does diverge — the weights grow without limit (see the math box). Real systems treat this with normalization (Oja's rule), lateral inhibition or competition between neurons. In the brain, stability is handled by homeostatic plasticity — a neuron adjusts its overall level of excitability and keeps the connections from running away.
Hebb is unsupervised learning. What exactly does the network "learn" if it is given no labels?
The correlation statistics of the inputs. As the link to PCA shows, a Hebbian neuron picks out the direction of maximum variance — the main structure in the data. At the level of assemblies this turns into associations: "what usually occurs together". There are no labels — there is co-occurrence.
Is this really how the brain learns?
Closer than backprop. Hebbian plasticity has a direct neural correlate — LTP/LTD (strengthening/weakening of synapses). The refined version, spike-timing-dependent plasticity, makes the rule causal: a connection strengthens if A fires shortly before B (not merely at the same time). So "together" actually means "A, then B".
What to read in the original
The book is mid-century neuropsychology; only one chapter, the one with the postulate, is ML-relevant. This write-up is enough. If you do dig in, look at the idea of cell assemblies: it is underratedly modern.