OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI
On the last day of its '12 Days of OpenAI', OpenAI previewed o3, which scored 75.7% on the ARC-AGI semi-private set (87.5% with high compute) — a benchmark on which earlier LLMs scored in single digits — and 25.2% on FrontierMath.
Key facts
- Announced 20 December 2024
- ARC-AGI-1 semi-private: 75.7% (high-efficiency), 87.5% (high-compute), verified by ARC Prize
- FrontierMath: 25.2% vs. under 2% for previous models, per OpenAI
- o3 and o4-mini released publicly on 16 April 2025, with full tool use
What happened
Only three months after o1, OpenAI showed that scaling RL and test-time compute produced another large leap in reasoning.
Why it matters
Convinced many observers that reasoning models were on a steep trajectory; ARC Prize called it a genuine step-change.
Changelog
- 2026-09-29: created
Related events
- OpenAI o1: reasoning models trained with reinforcement learning ★★★★★
- OpenAI o3 scores 25% on FrontierMath research-level maths benchmark, amid funding disclosure controversy ★★★
- OpenAI launches GPT-5 ★★★★★
Sources (2)
- officialOpenAI o3 Breakthrough High Score on ARC-AGI-Pub (ARC Prize)
- officialIntroducing OpenAI o3 and o4-mini (OpenAI)
id: 2024-12-20-openai-o3 · updated 2026-09-29 · open in the interactive timeline