Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. First Proof second batch: refereed benchmark finds AI…

First Proof second batch: refereed benchmark finds AI systems pass 7 of 10 unpublished research problems; ETH's ProofCouncil harness best with 6

★★★benchmarkFirst Proof FoundationETH ZurichOpenAIUCLAPrinceton Universityconfidence: high

On June 10, 2026 the First Proof Foundation (Abouzaid, Srivastava, Ward, Williams) released the formal second batch of its research-math benchmark. Ten solved but unpublished problems from working mathematicians were run by the organisers themselves on four public systems: ChatGPT 5.5 Pro and three academic harnesses. Each of the 39 solutions was refereed by 2–3 expert mathematicians. Seven problems got at least one passing solution (essentially flawless or minor revisions). The ETH Zurich/Aarhus ProofCouncil harness (System A, built on gpt-5.5-pro) passed 6. UCLA's harness and plain ChatGPT 5.5 Pro passed 5 each, the latter for $117 in total versus $3,186 and $4,799 for the harnesses. Princeton's Gemini-based harness passed 1, and no system made progress on the metric-geometry problem.

Key facts

What happened

After the informal first batch in February (2026-02-14-first-proof-challenge), the organisers set up the First Proof Foundation and ran batch 2 as a controlled benchmark. They tested only public systems, or harnesses over public models, ran every test themselves with full logs, and had each solution formally refereed. A smaller community-experimentation set was promised to follow.

Why it matters

It is the most carefully refereed measurement so far of AI on genuine research problems. By mid-2026, public systems produced publishable-quality solutions to most such problems, sometimes with new ideas. A plain chat model nearly matched harnesses that cost 30–40× more. The failure modes are the ones referees keep reporting: skipped hard steps and missing or false citations.

Changelog

  • 2026-10-05: created (leads run; from the ProofCouncil lead, arXiv 2607.09474)

Related events

  1. 'First Proof' challenge: AI solves about half of 10 unpublished research problems set by mathematicians ★★★

Sources (6)

id: 2026-06-10-first-proof-second-batch · updated 2026-10-05 · open in the interactive timeline