33 PPO
Context
Policy-gradient RL is unstable: one large update can wreck the policy, and recovering from that is hard. TRPO cured that with a hard constraint on the KL between the new and the old policy, but it was heavy going (second-order optimization). Schulman et al. simplify.
The idea and the mechanism
Instead of a hard constraint — a clipped surrogate objective. Take the ratio of the probabilities under the new and the old policy and "clip" it, so the update cannot run far from the old policy. That buys you several epochs of SGD on a single collected batch (more sample-efficient) with a plain first-order implementation.
reinforcement learning The clipped surrogate: a "pessimistic" bound
The ratio of the probability of an action under the new and the old policy:
The objective takes the minimum of the plain and the "clipped" version (Â is the estimated advantage of the action):
Take it apart by the sign of the advantage:
- Â > 0 (the action was good): we want to push r up, but the clip caps the gain once r > 1+ε — no incentive to run far.
- Â < 0 (bad): we want to push r down, and the clip puts a floor at 1−ε.
The minimum picks the more pessimistic of the two estimates → conservative steps without any explicit KL constraint. A simple first-order formula replaces TRPO's heavy optimization.
PyTorch The clipped PPO loss
import torch
def ppo_loss(logp, logp_old, adv, eps=0.2):
ratio = torch.exp(logp - logp_old) # π_new / π_old
clipped = torch.clamp(ratio, 1 - eps, 1 + eps)
return -torch.min(ratio * adv, clipped * adv).mean() # pessimistic (min) bound
Why it matters
Reliable, simple, few hyperparameters → the default deep RL algorithm. And, critically for this canon, PPO is the RL engine of RLHF: it is what InstructGPT/ChatGPT (#44) use to optimize an LLM against a reward model of human preferences. A caveat: the paper itself is about locomotion and Atari — the RLHF connection came later, as a downstream application.
Connections
DQN opened deep RL through the value function (learn Q). PPO is the other branch, policy-based (learn the policy directly), which works better for large or continuous policies. Both grow out of the success of deep RL; for LLMs it was the policy branch that turned out to matter.
In RLHF the language model is the policy, the reward model is the reward, and PPO optimizes the LM to maximize that reward under a KL penalty against the original model. A modest 2017 RL paper turned out to be a load-bearing part of what made LLMs into assistants.
RLHF on PPO is complicated (a separate reward model plus unstable RL). DPO shows you can reach the same result without RL — with a plain classification loss. PPO set the bar that DPO then cleared on simplicity.
Questions worth asking
What makes PPO better than TRPO, if the idea "do not go too far" is the same?
TRPO imposes a hard KL constraint and solves a second-order problem (conjugate gradient, Hessian-vector products) — complex and expensive. PPO gets similar stability out of plain clipping inside an ordinary first-order objective: easy to implement, works with standard SGD/Adam, and allows several epochs per batch. Simplicity won, not theoretical rigour.
Why clipping rather than a KL penalty in the loss?
The paper does include an adaptive-KL-penalty variant, but clipping is simpler and more robust: no penalty coefficient to pick and re-tune. The clip bounds the change in the policy directly, geometrically, with no fine-tuning of anything. In practice the clipped version became the standard precisely because it is robust out of the box.
A paper about robots — how did it become the basis of ChatGPT's alignment?
PPO is a general policy optimizer; it does not care that the policy happens to be a language model and the reward a score of human preferences. RLHF simply slotted the LLM into the role of the policy and the reward model into the role of the reward, and added a KL penalty against a reference. The 2017 authors had no idea their algorithm for making robots walk would be aligning conversational assistants — the classic story of a fundamental method finding an unexpected use.
What to read in the original
Short and practical — read it in full; the clipped objective is worth understanding exactly, since it underpins LLM alignment.