Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2022
  4. InstructGPT: RLHF aligns language models to follow…

InstructGPT: RLHF aligns language models to follow instructions

★★★★★researchOpenAIconfidence: high

OpenAI fine-tuned GPT-3 with reinforcement learning from human feedback (RLHF); labelers preferred outputs of the 1.3B InstructGPT over the 175B GPT-3, and the method became the recipe for ChatGPT.

Key facts

What happened

OpenAI made InstructGPT models the default in its API, showing that human-preference fine-tuning made models more helpful and truthful.

Why it matters

RLHF turned raw LLMs into usable assistants and underlies ChatGPT, Claude and nearly all chat models.

Changelog

  • 2026-09-29: created

Related events

  1. GPT-3 (175B) shows in-context few-shot learning ★★★★★
  2. OpenAI launches ChatGPT ★★★★★
  3. Anthropic introduces Constitutional AI (RLAIF) ★★★★

Sources (3)

id: 2022-01-27-instructgpt · updated 2026-09-29 · open in the interactive timeline