Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2024
  4. OpenAI o3 scores 25% on FrontierMath research-level maths…

OpenAI o3 scores 25% on FrontierMath research-level maths benchmark, amid funding disclosure controversy

★★★scienceOpenAIEpoch AIconfidence: high

Epoch AI's FrontierMath (Nov 2024) contains unpublished research-level problems on which models scored under 2%. On 20 Dec 2024 OpenAI claimed 25.2% for o3. It then emerged that OpenAI had funded the benchmark and had access to most problems. Released o3 scored lower in independent tests.

Key facts

Science result

Field
mathematics / benchmarks
Problem
FrontierMath: unpublished research-level problems with automatically checkable answers
Result
Score jump from <2% to a claimed 25.2% in about six weeks, later partly qualified by independent evaluation.
AI system
OpenAI o3
Human role
Autonomous answering; benchmark written by expert mathematicians
Verification
Company-reported; independent Epoch evaluation of released o3 was lower
Status
disputed
Why surprising
When FrontierMath launched, Fields medallists including Terence Tao said its problems would likely resist AI for years; a big jump came within weeks.

What happened

OpenAI previewed o3 with a headline FrontierMath score an order of magnitude above prior models. The benchmark's independence was then questioned when OpenAI's funding and access came to light.

Why it matters

It was the first sign that research-level maths was yielding to reasoning models, and an early lesson in benchmark governance and conflicts of interest.

Changelog

  • 2026-09-29: created

Related events

  1. OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI ★★★★★
  2. Google DeepMind's 'AI co-mathematician' sets FrontierMath Tier 4 record and helps resolve a Kourovka Notebook problem ★★★

Sources (3)

id: 2024-12-20-frontiermath-o3-25-percent · updated 2026-09-29 · open in the interactive timeline