InstructGPT: RLHF aligns language models to follow instructions
OpenAI fine-tuned GPT-3 with reinforcement learning from human feedback (RLHF); labelers preferred outputs of the 1.3B InstructGPT over the 175B GPT-3, and the method became the recipe for ChatGPT.
Key facts
- Announced 27 January 2022; paper arXiv 2203.02155
- Three steps: supervised fine-tuning, reward model, PPO optimization
- 1.3B InstructGPT outputs preferred over 175B GPT-3
- Built on 'Deep RL from Human Preferences' (Christiano et al., 2017, arXiv 1706.03741)
What happened
OpenAI made InstructGPT models the default in its API, showing that human-preference fine-tuning made models more helpful and truthful.
Why it matters
RLHF turned raw LLMs into usable assistants and underlies ChatGPT, Claude and nearly all chat models.
Changelog
- 2026-09-29: created
Related events
- GPT-3 (175B) shows in-context few-shot learning ★★★★★
- OpenAI launches ChatGPT ★★★★★
- Anthropic introduces Constitutional AI (RLAIF) ★★★★
Sources (3)
- officialAligning language models to follow instructions (OpenAI)
- paperTraining language models to follow instructions with human feedback (arXiv)
- paperDeep reinforcement learning from human preferences (arXiv)
id: 2022-01-27-instructgpt · updated 2026-09-29 · open in the interactive timeline