29 AlphaGo
Context
Go is the "grand challenge" of AI: ~10170 positions, brute force is out of the question, and a position is hard to score with a heuristic (unlike chess). DeepMind takes it on with deep networks plus search.
The idea and the mechanism
Three components. The policy network proposes likely moves (trained first on human games, then improved by self-play with RL). The value network scores a position (the probability of winning) without playing it out. Monte Carlo Tree Search is a guided search in which the policy narrows the BREADTH (which moves to look at) and the value the DEPTH (it cuts the playout short). The networks turn an impossible search into a manageable one.
search + RL How the networks steer the search tree (PUCT)
MCTS builds a tree; every edge (s,a) stores a visit count N, a mean value Q and a prior P from the policy network. Descending the tree means picking the move that maximizes "exploitation + exploration":
The first term pulls towards moves of high value; the second rewards rarely visited moves to which the policy network gave a high prior P. At a leaf, instead of a random playout to the end, the position is scored by the value network, and the score propagates back up, updating Q, N. So the policy trims the breadth and the value the depth, and out of 10170 what remains is a tree you can actually walk.
Python The selection step in MCTS (PUCT)
import math
def select(node, c=1.5):
total = sum(ch.N for ch in node.children)
def score(ch):
Q = ch.W / ch.N if ch.N else 0.0
U = c * ch.P * math.sqrt(total) / (1 + ch.N) # policy prior + exploration
return Q + U
return max(node.children, key=score) # descend the tree
Why it matters
A landmark for AI: it beat the professional Fan Hui 5:0 (in secret, 2015), then Lee Sedol 4:1 (2016, 200M+ viewers); "Move 37" was a creative move outside any human pattern. Its successor AlphaGo Zero (2017) learned with no human games at all, from pure self-play. It showed the power of "learning + search", which is now coming back in reasoning models (search/RL against verifiable rewards).
Connections
DQN proved that deep RL works for perception and control. AlphaGo takes the same "deep networks + RL" pairing and adds an explicit search (MCTS) — the next step in DeepMind's program after Atari.
The idea of "learning + search/verification" returns in reasoning LLMs: R1 teaches a model to reason through RL on verifiable rewards. The spirit is the same — improve the policy through interaction and checking — though the "search" has now unfolded into a chain of reasoning rather than a tree of moves.
AlphaGo combines RL with an explicit tree search; PPO is "pure" policy gradient with no search, optimizing the policy directly. Two schools of RL: with a model and search, and without. Both matter, and LLM alignment went down the second road (PPO in RLHF).
Questions worth asking
Why does brute-force search fail at Go when it worked at chess (Deep Blue)?
In chess the branching factor is ~35 and there is a good hand-written evaluation of a position (material, structure) — which let Deep Blue search deep with alpha-beta. In Go the branching is ~250 and the depth ~150 (hence the 10170), and even human experts find a position hard to formalize. So Go needed a learned evaluation (the value network) and learned move selection (the policy) — neither of which Deep Blue had.
"Move 37" — real creativity, or just good search?
It depends on your definition, but the effect is real: a move the system itself gave a ~1/10000 chance of being played by a human turned out to be strong. It came out of a value and a policy trained by self-play beyond human games — that is, AlphaGo explored regions people had pruned away as "wrong". A convincing example of superhuman novelty born of learning plus search rather than imitation.
AlphaGo Zero threw out the human games — why is that held to matter more than the original?
Because pure self-play from scratch beat the version trained on humans — meaning human data was a crutch, not a necessity. That gave a general template, "self-play RL + search", later generalized to chess and shogi (AlphaZero) and to learning a model of the environment (MuZero). Less human knowledge, more generality — DeepMind's signature trajectory.
What to read in the original
Read the key parts — the policy/value + MCTS pairing and the role of self-play. The MCTS math can be taken at the level of the idea (the math box); what matters more is seeing how learning and search reinforce each other.