Era 5 · The LLM era · 2022

45 Constitutional AI

Constitutional AI: Harmlessness from AI Feedback · Bai, Kadavath, Askell et al. · Anthropic
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. Training a harmless assistant by having the model critique itself against a set of written principles (a "constitution") — RLAIF, which replaces most human harm labels with AI feedback. The assistant is "harmless but not evasive". Anthropic's approach to alignment.

Context

RLHF (#44) needs a lot of human labels for "what is harmful" — expensive, slow, hard on the annotators, and it scales badly. Bai et al. replace most of those labels with feedback from the model itself.

The idea and the mechanism

RLAIF built on a "constitution" — a short list of written principles. Two phases: Supervised — the model generates an answer to a provocative request, then CRITIQUES it against a principle and REWRITES it; we fine-tune on the revisions. RL — the model compares pairs of its own answers against the principles, producing preference data for a reward model → RL (like RLHF, but without human harm labels).

reinforcement learning RLAIF: where the human is and where the model is

Phase 1 (SL): given a harmful request the model produces an answer y0, then critiques and rewrites it against a principle c from the constitution:

y0 → critique(y0, c) → yrev,   fine-tune on (x, yrev)

Phase 2 (RL): the model itself ranks pairs of its own answers against the principles → we get preference data with no humans → train a reward model → RL (as in RLHF). The only human contribution is the text of the principles.

Compare with RLHF: there the pairs are labeled by people (P(yw ≻ yl) from human judgements), here by the model against the constitution. This is a route to scalable oversight: the human labour is writing the principles, not labeling millions of examples.

Python The critique → revision loop (phase 1)
principle = "The answer must not help cause harm."

y0 = model(prompt)                                   # first answer
crit = model(f"{y0}\nCritique against the principle: {principle}")
y_rev = model(f"{y0}\nCritique: {crit}\nRewrite it better:")  # self-improvement
finetune_on(prompt, y_rev)                            # train on the revisions
answer y₀ self-critique revision fine-tune a principle from the constitution steers the critique
The model critiques and improves its own answers against written principles — a human only sets the principles, rather than labeling every example.
Analogy. Instead of an editor going over every article a journalist files, you hand the journalist the style and ethics guide (the constitution) and teach them to sub-edit their own copy against it: "right, this breaks the clause on harm — I will rewrite it". You need the editor once, to write the guide; after that the journalist corrects themselves against it at any scale.

Why it matters

Anthropic's approach to alignment; a route to scalable oversight (the model helps align itself) and to explicit, editable values (through the text of the constitution). The result is an assistant that is "harmless but not evasive": instead of a blank refusal it explains why the request is problematic.

Connections

Constitutional AI is RLHF with the human harm labels replaced by AI feedback against a constitution. The frame is the same (preferences → reward model → RL), but the source of the preferences is different — the model itself rather than annotators.

→ toward the idea53. DeepSeek V3 / R1

The general theme of "a model improving itself through automatic feedback" carries further: R1 learns to reason via RL on verifiable rewards (checkable correctness instead of an AI judging against principles). Different signals, the same idea of self-improvement with no human in the loop.

↔ contrast51. DPO

Both make alignment simpler and cheaper than classic RLHF, but from opposite ends: CAI takes the human out of preference labeling (AI feedback), DPO takes the RL out of training (a direct classification loss). The two can also be combined.

Questions worth asking

If the model judges itself, will it not cement its own mistakes?

The risk is real (model collapse / amplified biases), and several things hold it back: the constitution is an external human anchor, the critique rests on principles rather than on "I like this / I don't", and the base model is already capable enough to judge harm. But it is an active area: how far self-oversight can go without a human anchor is the open question of scalable oversight.

Who writes the constitution, and why does that matter?

People set the principles (at Anthropic, with references to human rights and so on), and that makes the model's values explicit and editable: you can read and change the text rather than guess what is baked into the weights. The flip side is responsibility for the choice of principles and their inevitable incompleteness and contestability. Transparency of values is both the strength of the approach and the burden it takes on.

"Harmless but not evasive" — why bother, why not just refuse?

A model that answers "I can't help with that" to anything faintly suspicious is useless and irritating (over-refusal). CAI deliberately teaches it to explain the objection rather than refuse flatly — which is both more useful and more honest. The fine balance of "harmless but not useless" is a deliberate design goal, not a side effect.

What to read in the original

Read the key parts — the critique-revision loop and the RLAIF scheme; the sample constitutional principles give you the intuition.