Era 1 · Origins · 1949

2 Hebbian learning

The Organization of Behavior (book) · Donald Hebb · Wiley
🟦 this write-up is enough~20–30 minoriginal ↗
The gist in 20 seconds. The first learning rule for neural networks: a connection strengthens when two neurons are active at the same time. Local, unsupervised, biologically plausible. Out of this single idea grow associative memory (Hopfield) and a link to principal component analysis.

Context

The McCulloch–Pitts neuron (1943) could compute but could not learn — its weights are fixed. The psychologist Donald Hebb goes looking for a biological mechanism that explains how experience changes the brain. His postulate will become the cornerstone of learning theory in neural networks.

The idea and the mechanism

Hebb's postulate: if neuron A systematically takes part in firing neuron B, the connection A→B strengthens. In its simplest form the weight change is proportional to the product of the activities:

Δwij = η · xi xj

The rule is local (it needs only the activities of the two connected neurons, no global error signal) and works unsupervised — the network soaks up the correlation statistics of its inputs. Hebb also introduced cell assemblies: groups of co-firing neurons that get "stitched together" into a stable pattern carrying a concept.

linear algebra Why pure Hebb is unstable — and how it connects to PCA

Take a linear neuron y = w·x and feed its output into the Hebb rule Δw = η y x. Averaging over the data:

E[Δw] = η · E[x(x·w)] = η · C w,   where C = E[x x⊤]

C is the covariance matrix of the inputs. This is power iteration: w grows along the leading eigenvector of C — the direction of maximum variance in the data. But the norm of w blows up without bound while it does so.

The cure is Oja's rule (1982), which adds a normalizing term:

Δw = η y (x − y w)

Now ‖w‖ → 1, and the neuron converges to the unit leading eigenvector — that is, it computes the first principal component (PCA). The conclusion: Hebb + normalization = learning a representation that picks out the main direction in the data.

NumPy Implementation: Oja's rule → the principal component
import numpy as np

def oja(X, eta=0.01, epochs=50):
    w = np.random.randn(X.shape[1]) * 0.01
    for _ in range(epochs):
        for x in X:
            y = w @ x                      # linear neuron
            w += eta * y * (x - y * w)     # Hebb + normalization (Oja)
    return w / np.linalg.norm(w)           # ≈ 1st principal component of X
A B connection strengthens A and B active together cell assembly:
Joint activity strengthens the A→B connection. Co-firing neurons stitch themselves into a cell assembly — a stable carrier of memory.
Analogy. Paths through a forest: the more often a route is walked, the more worn it becomes, and the more readily people take it again. Hebbian learning is this wearing-in of connections: the paths that signals travel together widen into highways. And Oja's rule is the forest warden who stops any one path from growing without limit.

Why it matters

This is the first rule by which a network changes itself from experience — the move from "a neuron as a gate" to "a neuron as a learning element". Out of it grow Hopfield's associative memory (its weights are set by the Hebb rule), self-organizing maps, and — on the biology side — long-term potentiation (LTP), the neural substrate of memory. The slogan "cells that fire together, wire together" is a late paraphrase (~1992), but it captures the point exactly.

Connections

The MP neuron gives you a unit of computation, but with fixed weights. Hebb adds a mechanism of plasticity to it — the connections change on their own. Together they are the first picture of a "learning network": computation (MP) plus learning (Hebb).

Hopfield takes the Hebb rule literally: the network's weights are set as a sum of outer products of the patterns to be stored. That turns an abstract postulate into a working model of associative memory with provable dynamics.

↔ contrast3. Perceptron

Two different spirits of learning. Hebb is unsupervised: it changes weights from the correlation of activities, without knowing the "right answer". The perceptron is supervised: it moves weights along an error signal relative to a label. That fork (unsupervised vs supervised) runs through the whole history of ML.

Questions worth asking

If the rule only strengthens connections, why doesn't the network "saturate" — every weight at maximum?

Pure Hebb really does diverge — the weights grow without limit (see the math box). Real systems treat this with normalization (Oja's rule), lateral inhibition or competition between neurons. In the brain, stability is handled by homeostatic plasticity — a neuron adjusts its overall level of excitability and keeps the connections from running away.

Hebb is unsupervised learning. What exactly does the network "learn" if it is given no labels?

The correlation statistics of the inputs. As the link to PCA shows, a Hebbian neuron picks out the direction of maximum variance — the main structure in the data. At the level of assemblies this turns into associations: "what usually occurs together". There are no labels — there is co-occurrence.

Is this really how the brain learns?

Closer than backprop. Hebbian plasticity has a direct neural correlate — LTP/LTD (strengthening/weakening of synapses). The refined version, spike-timing-dependent plasticity, makes the rule causal: a connection strengthens if A fires shortly before B (not merely at the same time). So "together" actually means "A, then B".

What to read in the original

The book is mid-century neuropsychology; only one chapter, the one with the postulate, is ML-relevant. This write-up is enough. If you do dig in, look at the idea of cell assemblies: it is underratedly modern.