OpenAI o1: reasoning models trained with reinforcement learning
OpenAI released o1-preview and o1-mini, models trained with large-scale RL to 'think' via a long private chain of thought before answering, yielding large gains in math, science and coding and introducing test-time compute scaling.
Key facts
- Announced 12 September 2024 (o1-preview, o1-mini); full o1 released 5 December 2024
- AIME 2024: o1 averaged 74% (single sample) vs 12% for GPT-4o, per OpenAI
- Exceeded PhD-level accuracy on GPQA Diamond science questions, per OpenAI
- Performance improved with both more RL training compute and more thinking time
- Codenamed 'Strawberry' in press reports
What happened
OpenAI introduced a new model series that spends variable inference-time compute reasoning before responding.
Why it matters
Opened the 'reasoning model' era and a new scaling axis (test-time compute); every major lab followed within months.
Changelog
- 2026-09-29: created
Related events
- Chain-of-thought prompting elicits reasoning in LLMs ★★★★
- OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI ★★★★★
- DeepSeek-R1: open-weights reasoning model rivals o1 and shakes markets ★★★★★
- AlphaGo Zero and AlphaZero master games through pure self-play ★★★★★
- OpenAI launches GPT-4o, a natively multimodal 'omni' model ★★★★
- Sam Altman publishes "The Intelligence Age": superintelligence possibly 'in a few thousand days' ★★★
Sources (2)
id: 2024-09-12-openai-o1 · updated 2026-09-29 · open in the interactive timeline