Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. 'First Proof' challenge: AI solves about half of 10…

'First Proof' challenge: AI solves about half of 10 unpublished research problems set by mathematicians

★★★scienceGoogle DeepMindOpenAIconfidence: high

Eleven mathematicians released 10 unpublished research-level problems on 5 Feb 2026 and answers on 14 Feb. DeepMind's Aletheia got 6/10 by majority expert assessment. OpenAI got at least 5 likely correct and retracted one claimed solution. Scientific American called the results 'mixed'.

Key facts

Science result

Field
mathematics / research-level problem solving
Problem
Ten previously unpublished lemmas and problems from working mathematicians
Result
Best AI systems solved roughly 5–6 of 10 fresh research problems, with some disputed gradings and one retraction.
AI system
Aletheia, OpenAI internal model
Human role
Autonomous attempts; human expert grading
Verification
Expert grading by the problem setters
Status
confirmed

What happened

Mathematicians created a contamination-proof test using problems from their own unpublished work, and AI labs submitted solutions within a week.

Why it matters

It gave a cleaner measure than olympiads of whether AI can do research maths: at the time, about half the time.

Changelog

  • 2026-09-29: created

Related events

  1. DeepMind's Aletheia agent and Gemini Deep Think report autonomous Erdős solutions and new physics and CS results ★★★★

Sources (3)

id: 2026-02-14-first-proof-challenge · updated 2026-09-29 · open in the interactive timeline