OpenProblemBench: 82 unresolved math and theoretical-physics problems
GPT-6 Astra has the highest model-judged solve rate (14%)
Confirmed
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 2 of 5
- Last verified
- 9 October 2026
Your AI and this story
- GPT-6 Astra161 days after its cutoff
- Claude Opus 5.5100 days after its cutoff
- Gemini 3.8 Flash191 days after its cutoff
- Grok 4.7130 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 100 days before it.
Key facts
- 82 problems; 7 solver configurations; 2,296 reviews; four evaluator models; outcome categories: solved, breakthrough partial, nontrivial partial, trivial partial, incomplete
- Solvers: GPT-6 Astra and GPT-5.6 Sol in Codex CLI; DeepSeek-V4.1-Flash in DeepSeek Harness; Kimi K3, GLM-5.3, Qwen3.8-Max and GLM-5.3-Flash in OpenCode; an extra Qwen run compares OpenCode with Claude Code
- GPT-6 Astra leads under every evaluator (15, 13, 9 and 9 solved); the union across configurations contains 13–16 problems judged solved by a fixed evaluator
- Qwen3.8-Max and Kimi K3 get solved or substantive-partial judgments in about 73–74% of reviews, GPT-5.6 Sol 63%; DeepSeek-V4.1-Flash is dominated by trivial partial outcomes (64.6%)
What happened
OpenProblemBench evaluates models on questions with no known answer. Grading is by other models, which check the stated obligations, quantifiers and decisive steps. The authors acknowledge that this measures judged progress, not expert-verified solutions.
Why it matters
It adds a benchmark of open research problems, after FrontierMath-style tests with known answers. It ranks the AI systems that are producing real conjecture results in 2026, with GPT-6 Astra clearly ahead and Chinese open models close to each other.
Sources
1 source from 1 site. Numbers match the chips in the text.
1 source: 1 primary
Primary
Changes
- Filed from the arXiv PDF