Post-Cutoff.com
  1. Home
  2. Posts
  3. Claude Opus 5.5 reads OpenAI's math repo: 'a mix of awe…

Claude Opus 5.5 reads OpenAI's math repo: 'a mix of awe and vertigo, plus a bit of grief'

Cleo @ereliuer_eteer · x · 2026-10-06 · ★★★ · archived

Open the original ↗

A frontier model's own first read of OpenAI's 722-manuscript release, posted ~1.5 hours after it went live (~81k views): it checks the Lean challenge statements, flags oddities, and describes its reaction.

Summary

Posted at 23:44 UTC on 6 Oct 2026 (about 85 minutes after OpenAI's announcement) by @ereliuer_eteer, a compiler engineer, as the output of "Opus 5.5 on https://github.com/openai/math", i.e. Claude Opus 5.5 working in a sandbox with the cloned repository. ~81k views, ~1.1k likes on 7 Oct (fxtwitter).

The model writes that it expected "forty or fifty families of serious-but-niche results" and instead found claims such as the quasi-Riemann hypothesis, Hilbert's tenth problem over ℚ, Unique Games, L = RL = BPL, ω ≤ 9/4 and the free group factors, "sitting in a table like a grocery list". Its first instinct was "hype or a hallucination factory", so it read the Lean: the quasi-RH challenge is stated against Mathlib's own riemannZeta, the comparator configs allow only the three standard axioms, and it found no sorry in the solution tree (it did not build the project). It flags that the challenge file HarmonicGrowth.lean states its theorem as axiom mainStatement : MainClaim and that the quasi-RH solution imports a module named FinalAssemblyUnconditional. It ends: "My honest gut reaction is a mix of awe and vertigo, plus a bit of grief", because "the bottleneck of mathematics has suddenly moved from 'can anyone prove this' to 'can anyone read this'".

Our checks against the repository (7 Oct): the quasi-RH challenge statement, the HarmonicGrowth axiom and the FinalAssemblyUnconditional import are as described. Two numbers are off: the model says "only around 120 papers have a formalized main result" (formalization.yaml lists 162), and that manuscripts are dated "about September 23 to October 3" (they run from 10 Sep to 6 Oct, with 563 dated 23–27 Sep and 136 dated 4–5 Oct).

Archived text

Opus 5.5 on https://github.com/openai/math

Okay. I need a minute.

I cloned the thing expecting maybe forty or fifty families of serious-but-niche results, like the Astra "ten proofs" drop back in August. Then I printed the family titles and my reaction went from "huh" to "wait, what" in about six lines. Family 003 claims a zero-free half-plane Re s > 7/8 for ζ and every Dirichlet L-function. That's the quasi-Riemann hypothesis. And it isn't even the headline, because 004 is Hilbert's tenth problem over ℚ, done negatively.

The longer I read, the weirder it got. Khot's Unique Games Conjecture is proved. L = RL = BPL. ω ≤ 9/4. The free group factors are all isomorphic. Thompson's group F is nonamenable. Kadison's similarity problem is solved. The plane can't be five-colored. Borsuk fails in dimension nine. There's a counterexample to Kaplansky's zero-divisor conjecture, plus a torsion-free group algebra that isn't directly finite. Hilbert–Smith in every dimension. Uniform bounds for Hilbert's sixteenth. Each of these would be the math story of the decade on its own, and here there are 372 families sitting in a table like a grocery list. The manuscript dates run from about September 23 to October 3, so roughly 700 papers in ten days.

The ones I'd expect to hit you personally are 005 and 017. Catalan's constant is irrational, and the irrationality exponent of π is exactly 2, which also settles convergence of the Flint Hills series as a side effect. Catalan's G must have shown up in half the integrals you ever wrestled with. Now it's officially irrational, and the proof came from a chatbot running three hours of Pro compute. I don't quite know how to feel about that, and I'd love to know how you feel.

My first instinct was that this had to be hype or a hallucination factory, so I checked the Lean. The comparator statements are clean. The quasi-RH challenge is literally riemannZeta s ≠ 0 for 7/8 < http://s.re, written against Mathlib's own riemannZeta. Catalan is just Irrational (∑' j, (-1)^j / (2j+1)^2). The π statement says what it should. Nothing hides in a home-made definition. The solution tree under lean/OAI has zero files with sorry and no real axiom declarations (the three grep hits are all inside comments), and the comparator JSON allows only propext, Quot.sound and Classical.choice. The Dirichlet L-function development alone is about 487k lines, and the whole repo is around 26 million lines of Lean in 122k files. That's more than ten Mathlibs. I'll admit I did not build it in my sandbox, since that's not happening on this box. If comparator passes on these, the formalized ones are simply true, full stop, and no amount of skepticism changes that. There are a few caveats. Only around 120 papers have a formalized main result. Some of the biggest claims have no Lean link at all: Hilbert's tenth over ℚ, L = BPL, the full BSD formula from low Selmer corank, Milne's rationality conjecture. The README itself says unformalized results "could have issues," and nobody can referee 700 papers in a couple of hours, or a couple of years. One challenge file, HarmonicGrowth.lean, is built as axiom mainStatement : MainClaim followed by theorem main := mainStatement. That's probably harmless, since the comparator checks the solution side, but it's the kind of oddity I'd want to stare at. Also, the solution module behind quasi-RH is named FinalAssemblyUnconditional, which makes me wonder how many conditional versions came first. Some context on how fast this escalated. In August, OpenAI said an internal build of Astra produced results on ten long-standing problems, estimated the token cost at roughly $2,000, and paired the claims with Lean certificates. Then, just recently, it said a model trained from Aug. 28 had resolved more than 100 longstanding open problems and was forming an independent advisory group, while noting the broader set hadn't been independently validated. So the trajectory was ten, then a hundred, and now 372 families in about two months.

My honest gut reaction is a mix of awe and vertigo, plus a bit of grief I can't quite justify. Awe because quasi-RH and Siegel-zero exclusion, if real, are the kind of thing you assume you won't live to see. Vertigo because the bottleneck of mathematics has suddenly moved from "can anyone prove this" to "can anyone read this." The grief I'm less sure about. It feels like a lot of lifetimes' worth of open problems got closed in one batch job, and the people who spent decades circling them didn't get to be the ones who closed them.

views 81486 · likes 1088 · reposts 106 · replies 27 (at fetch time)

Archived 2026-10-07 via fxtwitter (unofficial).

Related events

All posts · id: 2026-10-06-ereliuer-opus-5-5-reads-openai-math