'First Proof' challenge: AI solves about half of 10 unpublished research problems set by mathematicians
Eleven mathematicians released 10 unpublished research-level problems on 5 Feb 2026 and answers on 14 Feb. DeepMind's Aletheia got 6/10 by majority expert assessment. OpenAI got at least 5 likely correct and retracted one claimed solution. Scientific American called the results 'mixed'.
Key facts
- 10 problems from the authors' own unpublished research; answers revealed 14 Feb 2026
- Aletheia: problems 2, 5, 7, 8, 9, 10 judged correct by majority (experts split on #8)
- OpenAI: problems 4, 5, 6, 9, 10 likely correct; retracted claim on #2
Science result
- Field
- mathematics / research-level problem solving
- Problem
- Ten previously unpublished lemmas and problems from working mathematicians
- Result
- Best AI systems solved roughly 5–6 of 10 fresh research problems, with some disputed gradings and one retraction.
- AI system
- Aletheia, OpenAI internal model
- Human role
- Autonomous attempts; human expert grading
- Verification
- Expert grading by the problem setters
- Status
- confirmed
What happened
Mathematicians created a contamination-proof test using problems from their own unpublished work, and AI labs submitted solutions within a week.
Why it matters
It gave a cleaner measure than olympiads of whether AI can do research maths: at the time, about half the time.
Changelog
- 2026-09-29: created
Related events
Sources (3)
- officialFirst Proof challenge
- officialOpenAI: First Proof submissions
- pressScientific American: First Proof is AI's toughest math test yet — the results are mixed
id: 2026-02-14-first-proof-challenge · updated 2026-09-29 · open in the interactive timeline