First Proof second batch: refereed benchmark finds AI systems pass 7 of 10 unpublished research problems; ETH's ProofCouncil harness best with 6
On June 10, 2026 the First Proof Foundation (Abouzaid, Srivastava, Ward, Williams) released the formal second batch of its research-math benchmark. Ten solved but unpublished problems from working mathematicians were run by the organisers themselves on four public systems: ChatGPT 5.5 Pro and three academic harnesses. Each of the 39 solutions was refereed by 2–3 expert mathematicians. Seven problems got at least one passing solution (essentially flawless or minor revisions). The ETH Zurich/Aarhus ProofCouncil harness (System A, built on gpt-5.5-pro) passed 6. UCLA's harness and plain ChatGPT 5.5 Pro passed 5 each, the latter for $117 in total versus $3,186 and $4,799 for the harnesses. Princeton's Gemini-based harness passed 1, and no system made progress on the metric-geometry problem.
Key facts
- Systems: A = IMProofBench ProofCouncil (ETH Zurich/Aarhus; base gpt-5.5-pro, limited use of gpt-5.5, gemini-3.1-pro-preview, claude-opus-4-7); B = UCLA Moonshot harness (gpt-5.5-pro); C = OpenAI ChatGPT 5.5 Pro; D = Princeton Momus (gemini-3.1-pro-preview)
- Passing grades per system (our tally of the report's editorial decisions): A 6 of 9 graded (P1, P2, P3, P5, P7, P9; no P6 solution due to a technical issue), B 5 (P1, P2, P6, P7, P9), C 5 (P1, P2, P6, P7, P9), D 1 (P7). Problems 1, 2, 3, 5, 6, 7, 9 had at least one pass; P8 and P10 only 'major revisions'; P4 (metric geometry) no progress
- Problem 5 (stochastic PDE): System A's solution rated essentially flawless, with a novel approach different from the human solution that 'impressed the referees'; AI arguments also differed from the human ones on Problems 3 and 9
- Cost and time: A $3,186 / 22.9 h; B $4,799 / 23.1 h; C $117 / ~5.8 h; D ~$1,014 (imputed) / 7.8 h. The report: harnesses can improve on their base model, 'However, this improvement comes at a substantial financial cost'
- Referee themes: routine steps handled meticulously while the hardest steps were glossed as 'standard arguments'; citations to papers that do not contain the claimed results; several Problem 2 solutions copied the author's earlier paper's terminology and labels without citing it, which for a human 'would have been flagged for plagiarism'
- Process: problems collected March–May; systems fixed by May 28; grading with 30 expert mathematicians at Harvard CMSA June 4–5; report and webinar June 10; arXiv 2606.18119 (June 16)
- ProofCouncil paper (arXiv 2607.09474, July 10; Schmitt, Gehrunger, Dekoninck, Bérczi, Kreitner, Price, Holmes): author–critic agent, open-sourced as eth-sri/proof-council; on 30 researcher-submitted open problems, 5 judged completely correct, 2 promising pending verification, 8 with useful partial progress
What happened
After the informal first batch in February (2026-02-14-first-proof-challenge), the organisers set up the First Proof Foundation and ran batch 2 as a controlled benchmark. They tested only public systems, or harnesses over public models, ran every test themselves with full logs, and had each solution formally refereed. A smaller community-experimentation set was promised to follow.
Why it matters
It is the most carefully refereed measurement so far of AI on genuine research problems. By mid-2026, public systems produced publishable-quality solutions to most such problems, sometimes with new ideas. A plain chat model nearly matched harnesses that cost 30–40× more. The failure modes are the ones referees keep reporting: skipped hard steps and missing or false citations.
Changelog
- 2026-10-05: created (leads run; from the ProofCouncil lead, arXiv 2607.09474)
Related events
Sources (6)
- paperFirst Proof Second Batch report (June 10, 2026)
- paperarXiv 2606.18119: First Proof Second Batch
- officialFirst Proof Project
- pressHarvard Math: First Proof's second batch of math problems test AI
- pressNature news (June 12, 2026): Humans outperform AI at this highly rigorous mathematics test (PDF copy)
- paperarXiv 2607.09474: ProofCouncil, an LLM agent for solving open mathematical problems
id: 2026-06-10-first-proof-second-batch · updated 2026-10-05 · open in the interactive timeline