Era 2 · Foundations · 2003

13 Neural language model

A Neural Probabilistic Language Model · Bengio, Ducharme, Vincent, Jauvin · JMLR
🟧 read selectively~1.5–2 horiginal ↗
The gist in 20 seconds. A language model in which every word is a learned vector (an embedding) and a neural network predicts the next word. Similar words end up with nearby vectors → the model generalizes to combinations it has never seen. The direct ancestor of word2vec and of every LLM.

Context

Classical language models are counting n-grams: they estimate the probability of a word from the frequencies of the words before it. The trouble is the "curse of dimensionality": almost any test phrase has never occurred verbatim, and there is no notion of words being similar. Bengio (2003) proposes a neural LM.

The idea and the mechanism

Every word is assigned a learned vector ew ∈ ℝd. The model takes the embeddings of the n−1 previous words, concatenates them, runs them through a hidden layer and a softmax → a distribution over the next word. The embeddings are learned jointly with the task. Words used in similar ways end up with nearby vectors → the model generalizes: having learned "the cat sits on …", it predicts "the dog lies on …" better.

probability The curse of dimensionality, and why embeddings lift it

A counting n-gram stores a probability for every possible context of n words. With a vocabulary of V, the number of contexts is:

Vn  (for V=10⁵, n=4 that is 10²⁰)

Most contexts never occur in the data at all → their probability is zero and there is no generalization. A neural LM replaces the table with a function:

P(wt = j | context) = softmax(Uh)j = eojΣk eok

The parameter count is now ≈ V·d (the embeddings) plus the network weights — linear in V, not exponential. And crucially: embeddings share statistical strength between similar words — nearby vectors give nearby predictions, so an unseen combination gets a sensible probability.

NumPy The forward pass of a neural language model
import numpy as np

def neural_lm(ctx, E, W, b, U):
    x = E[ctx].reshape(-1)            # embeddings of the n−1 words, concatenated
    h = np.tanh(W @ x + b)           # hidden layer
    logits = U @ h                   # scores over the whole vocabulary
    p = np.exp(logits - logits.max())
    return p / p.sum()               # softmax → P(next word)
"cat" "sits" "on" emb.ℝᵈ tanh softmaxover the vocabulary P(wₜ)
Words → embeddings (shared across all contexts) → hidden layer → softmax over the vocabulary. Similar words — nearby vectors — nearby predictions.
Analogy. A counting n-gram is a huge phrasebook in which every phrase is spelled out literally; no exact phrase, no answer — the model is mute. A neural LM is a person who understands the meaning of words: on hearing an unfamiliar combination, they substitute words of similar meaning and guess how it continues. The embeddings are that "meaning", moved into geometry.

Why it matters

The direct intellectual ancestor of word2vec/GloVe (where embeddings became an end in themselves) and of every modern LLM: the scheme is the same — tokens → embeddings → context → softmax, only now the context is built by a Transformer rather than a small MLP. The limitation of the era was cost (a large softmax, the hardware of 2003).

Connections

→ leads to17. Word2Vec

Word2vec takes the by-product of this model — the embeddings — and makes them the main goal, throwing out the expensive hidden layer for speed. Word vectors go from being "an implementation detail of a language model" to a standalone tool used everywhere.

→ leads to32. Transformer

The basic LLM scheme is laid down here: represent words as vectors and predict the next one through a softmax. The Transformer radically strengthens the "context" part (self-attention instead of a fixed window of n−1 words), but the skeleton is the same.

↔ contrast10. SVM

Two views on "how to represent data": an SVM takes a fixed (kernel) feature space, specified by hand; a neural language model learns its representation (the embeddings) for the task. The "fixed features vs learned features" argument shows up here too.

Questions worth asking

Where does "meaning" in the embeddings come from at all, if the model only predicts the next word?

From the distributional hypothesis: "you shall know a word by the company it keeps". To predict context well, the model is forced to place words with similar surroundings close together in the space — and so the geometry starts to encode semantics. Nobody labeled "meaning"; it emerges as a side effect of the prediction task.

A softmax over the whole vocabulary is expensive. How do people deal with that?

It is the main bottleneck. The cure is approximation: hierarchical softmax (a tree instead of flat normalization), negative sampling and noise-contrastive estimation (learn to tell the real word from random ones, without normalizing over the whole vocabulary — this is what word2vec did), sampled softmax. All of them sidestep the sum over V words.

If the 2003 idea is so powerful, why did the LLM revolution only happen ~17 years later?

Three things were missing: data (huge corpora), compute (GPUs/TPUs) and an architecture for context (the Transformer instead of a fixed window). Bengio wrote down the recipe, but it scaled badly on the hardware of 2003. The history of DL is largely about old, correct ideas waiting for data and compute.

What to read in the original

Read the essentials — the idea of a distributed representation and the architecture; this is the intellectual ancestor of embeddings and of LLMs, and it is worth seeing it in its original form.