Anthropic introduces Constitutional AI (RLAIF)
Anthropic's Constitutional AI trained a harmless-but-helpful assistant using AI feedback guided by a written set of principles (a 'constitution') instead of human harm labels.
Key facts
- arXiv 2212.08073 'Constitutional AI: Harmlessness from AI Feedback' (December 2022)
- Two phases: supervised self-critique and revision, then RL from AI feedback (RLAIF)
- Used in training Anthropic's Claude models
- Anthropic published Claude's constitution in May 2023
What happened
The model critiqued and revised its own outputs according to principles, and a preference model trained on AI judgments then guided RL.
Why it matters
Showed alignment could scale with AI supervision, making values explicit and auditable; RLAIF is now widespread.
Changelog
- 2026-09-29: created
Related events
- InstructGPT: RLHF aligns language models to follow instructions ★★★★★
- Anthropic releases Claude ★★★★
- Anthropic launches with a focus on AI safety ★★★
Sources (2)
id: 2022-12-15-constitutional-ai · updated 2026-09-29 · open in the interactive timeline