13 Neural language model
Context
Classical language models are counting n-grams: they estimate the probability of a word from the frequencies of the words before it. The trouble is the "curse of dimensionality": almost any test phrase has never occurred verbatim, and there is no notion of words being similar. Bengio (2003) proposes a neural LM.
The idea and the mechanism
Every word is assigned a learned vector ew ∈ ℝd. The model takes the embeddings of the n−1 previous words, concatenates them, runs them through a hidden layer and a softmax → a distribution over the next word. The embeddings are learned jointly with the task. Words used in similar ways end up with nearby vectors → the model generalizes: having learned "the cat sits on …", it predicts "the dog lies on …" better.
probability The curse of dimensionality, and why embeddings lift it
A counting n-gram stores a probability for every possible context of n words. With a vocabulary of V, the number of contexts is:
Most contexts never occur in the data at all → their probability is zero and there is no generalization. A neural LM replaces the table with a function:
The parameter count is now ≈ V·d (the embeddings) plus the network weights — linear in V, not exponential. And crucially: embeddings share statistical strength between similar words — nearby vectors give nearby predictions, so an unseen combination gets a sensible probability.
NumPy The forward pass of a neural language model
import numpy as np
def neural_lm(ctx, E, W, b, U):
x = E[ctx].reshape(-1) # embeddings of the n−1 words, concatenated
h = np.tanh(W @ x + b) # hidden layer
logits = U @ h # scores over the whole vocabulary
p = np.exp(logits - logits.max())
return p / p.sum() # softmax → P(next word)
Why it matters
The direct intellectual ancestor of word2vec/GloVe (where embeddings became an end in themselves) and of every modern LLM: the scheme is the same — tokens → embeddings → context → softmax, only now the context is built by a Transformer rather than a small MLP. The limitation of the era was cost (a large softmax, the hardware of 2003).
Connections
Word2vec takes the by-product of this model — the embeddings — and makes them the main goal, throwing out the expensive hidden layer for speed. Word vectors go from being "an implementation detail of a language model" to a standalone tool used everywhere.
The basic LLM scheme is laid down here: represent words as vectors and predict the next one through a softmax. The Transformer radically strengthens the "context" part (self-attention instead of a fixed window of n−1 words), but the skeleton is the same.
Two views on "how to represent data": an SVM takes a fixed (kernel) feature space, specified by hand; a neural language model learns its representation (the embeddings) for the task. The "fixed features vs learned features" argument shows up here too.
Questions worth asking
Where does "meaning" in the embeddings come from at all, if the model only predicts the next word?
From the distributional hypothesis: "you shall know a word by the company it keeps". To predict context well, the model is forced to place words with similar surroundings close together in the space — and so the geometry starts to encode semantics. Nobody labeled "meaning"; it emerges as a side effect of the prediction task.
A softmax over the whole vocabulary is expensive. How do people deal with that?
It is the main bottleneck. The cure is approximation: hierarchical softmax (a tree instead of flat normalization), negative sampling and noise-contrastive estimation (learn to tell the real word from random ones, without normalizing over the whole vocabulary — this is what word2vec did), sampled softmax. All of them sidestep the sum over V words.
If the 2003 idea is so powerful, why did the LLM revolution only happen ~17 years later?
Three things were missing: data (huge corpora), compute (GPUs/TPUs) and an architecture for context (the Transformer instead of a fixed window). Bengio wrote down the recipe, but it scaled badly on the hardware of 2003. The history of DL is largely about old, correct ideas waiting for data and compute.
What to read in the original
Read the essentials — the idea of a distributed representation and the architecture; this is the intellectual ancestor of embeddings and of LLMs, and it is worth seeing it in its original form.