Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2024
  4. OpenAI o1: reasoning models trained with reinforcement…

OpenAI o1: reasoning models trained with reinforcement learning

★★★★★model-releaseOpenAIconfidence: high

OpenAI released o1-preview and o1-mini, models trained with large-scale RL to 'think' via a long private chain of thought before answering, yielding large gains in math, science and coding and introducing test-time compute scaling.

Key facts

What happened

OpenAI introduced a new model series that spends variable inference-time compute reasoning before responding.

Why it matters

Opened the 'reasoning model' era and a new scaling axis (test-time compute); every major lab followed within months.

Changelog

  • 2026-09-29: created

Related events

  1. Chain-of-thought prompting elicits reasoning in LLMs ★★★★
  2. OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI ★★★★★
  3. DeepSeek-R1: open-weights reasoning model rivals o1 and shakes markets ★★★★★
  4. AlphaGo Zero and AlphaZero master games through pure self-play ★★★★★
  5. OpenAI launches GPT-4o, a natively multimodal 'omni' model ★★★★
  6. Sam Altman publishes "The Intelligence Age": superintelligence possibly 'in a few thousand days' ★★★

Sources (2)

id: 2024-09-12-openai-o1 · updated 2026-09-29 · open in the interactive timeline