23 Adam
Context
SGD needs the learning rate hand-tuned carefully and behaves badly when gradients are sparse or live on very different scales. Kingma and Ba give us a near-universal optimizer.
The idea and the mechanism
For every parameter we keep two running averages of the gradient: the first moment m (momentum-like — a smoothed direction) and the second moment v (RMSProp-like — a smoothed square, i.e. a scale). Dividing the step by √v gives per-parameter adaptive steps: where gradients are often large the step is smaller, where they are rare it is larger.
optimization The moments, and why bias correction is needed
Running (exponential) averages of the first and second moment of the gradient:
The start-up problem. We initialize m0 = v0 = 0, so on the first steps the estimates are biased low. Unrolling the recursion you can show E[mt] = (1 − β1t) E[g] — hence the correction by division:
The final step is the moment divided by the square root of the scale:
Dividing by √v̂ normalizes each parameter to its own gradient scale — hence the stability with barely any learning-rate search.
Python One Adam step
def adam_step(w, g, m, v, t, lr=1e-3, b1=0.9, b2=0.999, eps=1e-8):
m = b1 * m + (1 - b1) * g # 1st moment (momentum)
v = b2 * v + (1 - b2) * g * g # 2nd moment (scale)
mh = m / (1 - b1 ** t) # bias correction
vh = v / (1 - b2 ** t)
w -= lr * mh / (vh ** 0.5 + eps) # adaptive per-parameter step
return w, m, v
Why it matters
"Set it and it works", with almost no tuning → the default optimizer of deep learning, especially for transformers (usually AdamW — with weight decay done properly). One wrinkle: the original convergence proof turned out to be WRONG (a counterexample in AMSGrad, ICLR-2018), and sometimes a well-tuned SGD generalizes better.
Connections
A clean division of labour: backprop answers "which way" (it computes the gradient), Adam answers "with what step" (how to move along it). One gives the direction, the other the scale and the inertia; together they are what training a network is.
Practically every large model is trained with a variant of Adam (AdamW). Without a stable adaptive optimizer, training transformers on enormous data would be far more temperamental — Adam is the load-bearing part nobody notices.
Two sides of the same question, "how do you train a deep network": Adam improves the optimizer, ResNet the architecture (skip connections). Sometimes a good architecture matters more than a clever optimizer, sometimes the other way round — and both lines developed in parallel.
Questions worth asking
If the convergence proof was wrong, why does Adam work anyway?
Because "works in practice" and "provably converges in the worst case" are different things. AMSGrad exhibited a counterexample where Adam does not converge, and offered a fix — but on real deep-learning problems the original Adam behaves beautifully. Theory was catching up with practice; a familiar plot in ML, where empirical success runs ahead of guarantees.
People say SGD generalizes better than Adam — so why does everyone use Adam?
Adam is faster and steadier at the start and needs less learning-rate search — critical for huge models, where a single run is expensive. On some problems (classical vision) a carefully tuned SGD+momentum generalizes slightly better, but that takes fiddling. For transformers Adam/AdamW is all but the only option — SGD trains them badly.
Why AdamW if Adam exists — what is the difference?
How weight decay is applied. In plain Adam the L2 regularization goes into the gradient and is then divided by √v̂ — it gets scaled unevenly and does not work as intended. AdamW decouples weight decay from the adaptive step (applying it directly to the weights), and the regularization behaves correctly again. A small fix with a noticeable effect on generalization.
What to read in the original
This write-up plus the algorithm box from the paper is enough. If you want to dig: the derivation of the bias correction (the math box) and the discussion of convergence in AMSGrad.