18 DQN (Atari)
Context
RL could train agents in small or hand-crafted state spaces. DeepMind learns to play Atari straight from PIXELS, with no hand-designed features — by joining Q-learning to a deep network.
The idea and the mechanism
The value function Q(s, a) (the expected future reward for taking action a in state s) is approximated by a convolutional network: the input is the last few screen frames, the output is a Q value per action, and the agent takes argmax Q. Two tricks, without which training fell apart: experience replay (a buffer of transitions, sampled at random → breaks the correlation between neighbouring frames) and a target network (a frozen copy used to compute the targets → stability).
reinforcement learning The Bellman equation and why you need a target network
The optimal Q function satisfies the Bellman equation: value now = reward + the discounted best value from there on:
DQN approximates Q* with a network Qθ by minimizing the temporal-difference error — the gap between the left and right sides:
The subtlety: Q appears both in the target and in the prediction. If the target is computed by the same network you are training, it runs away at every step — training chases a moving target and diverges. The target network Qθ⁻ — a frozen copy, updated rarely — pins the target down and stabilizes the process.
PyTorch The DQN TD loss with a target network
import torch
def dqn_loss(batch, Q, Q_target, gamma=0.99):
s, a, r, s2, done = batch
q = Q(s).gather(1, a) # Q(s, a)
with torch.no_grad():
target = r + gamma * Q_target(s2).max(1).values * (1 - done)
return ((q.squeeze() - target) ** 2).mean() # temporal-difference error
Why it matters
It launched deep RL: one network with one set of hyperparameters learned 7 games from raw pixels, several of them at or above human level. A straight line runs from here to AlphaGo (#29) and to the RL half of RLHF (#44). It showed that DL + RL give you agents that learn perception and control end-to-end.
Connections
The agent's "eyes" are a convolutional network reading the screen. DQN is CNN perception with a Q-learning objective bolted on top; without mature CNNs, learning from pixels would have been impossible.
DQN proved that deep RL works for perception and control. AlphaGo takes the same "deep networks + RL" pairing and adds search (MCTS) to crack Go — the next rung of DeepMind's program.
DQN is value-based RL (learn Q, the action is the argmax). PPO is policy-based (learn the policy directly), which is steadier for continuous actions and large policies. It is policy gradients (PPO) that will end up powering RLHF, but both branches grow out of that same success of deep RL, the one DQN kicked off.
Questions worth asking
Why is RL so unstable compared with supervised learning?
Three troubles meet (the "deadly triad"): (1) bootstrapping — the target comes from the network's own estimates; (2) function approximation; (3) off-policy data. On top of that the data distribution is non-stationary — the policy changes, and so do the states you encounter. Replay and the target network damp part of the problem, but RL stays more temperamental than training against fixed labels.
DQN plays from pixels — is that a step towards general AI?
Yes and no. The generality is impressive: one architecture handled many different games with no tuning. But it is narrow: the model learns each game from scratch, transfers no skill, and is wildly data-hungry (millions of frames per game). It is an important milestone in the deep RL program, but a "general" agent is far off — and sample efficiency remains RL's sore point.
Q-learning overestimates values — that is a known problem, what does it have to do with DQN?
The max operator in the target systematically inflates Q (take a maximum over noisy estimates → an upward bias). DQN suffers from it; the cure is Double DQN: pick the action with one network and evaluate it with another (the target), decoupling the bias in selection from the bias in evaluation. A nice example of how a fine statistical detail (max is biased) spoils training.
What to read in the original
It is short — read it in full; pay attention to experience replay and the target network: these are general stabilization patterns that resurface across all of deep RL.