OpenAI o3 scores 25% on FrontierMath research-level maths benchmark, amid funding disclosure controversy
Epoch AI's FrontierMath (Nov 2024) contains unpublished research-level problems on which models scored under 2%. On 20 Dec 2024 OpenAI claimed 25.2% for o3. It then emerged that OpenAI had funded the benchmark and had access to most problems. Released o3 scored lower in independent tests.
Key facts
- FrontierMath paper v1: 7 Nov 2024; prior models <2%
- o3 claimed 25.2% (aggressive test-time compute setting) on 20 Dec 2024
- OpenAI's funding and data access disclosed only in paper v5 (20 Dec 2024); Epoch said it should have been more transparent
- Later records: Gemini 3 Pro 38% (Tiers 1-3) and 19% (Tier 4) in Nov 2025; GPT-5.2 Pro 31% on Tier 4 in Jan 2026
Science result
- Field
- mathematics / benchmarks
- Problem
- FrontierMath: unpublished research-level problems with automatically checkable answers
- Result
- Score jump from <2% to a claimed 25.2% in about six weeks, later partly qualified by independent evaluation.
- AI system
- OpenAI o3
- Human role
- Autonomous answering; benchmark written by expert mathematicians
- Verification
- Company-reported; independent Epoch evaluation of released o3 was lower
- Status
- disputed
- Why surprising
- When FrontierMath launched, Fields medallists including Terence Tao said its problems would likely resist AI for years; a big jump came within weeks.
What happened
OpenAI previewed o3 with a headline FrontierMath score an order of magnitude above prior models. The benchmark's independence was then questioned when OpenAI's funding and access came to light.
Why it matters
It was the first sign that research-level maths was yielding to reasoning models, and an early lesson in benchmark governance and conflicts of interest.
Changelog
- 2026-09-29: created
Related events
- OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI ★★★★★
- Google DeepMind's 'AI co-mathematician' sets FrontierMath Tier 4 record and helps resolve a Kourovka Notebook problem ★★★
Sources (3)
- officialEpoch AI: OpenAI and FrontierMath
- pressThe Decoder: OpenAI quietly funded independent math benchmark
- pressTechRepublic: independent FrontierMath score for o3
id: 2024-12-20-frontiermath-o3-25-percent · updated 2026-09-29 · open in the interactive timeline