--- id: "2026-10-08-openproblembench-82-open-problems" url: "https://postcutoff.com/e/2026-10-08-openproblembench-82-open-problems/" as_of: "2026-10-09T19:24:00+02:00" date: "2026-10-08" date_precision: day category: benchmark importance: 2 confidence: high status: [Confirmed] sources: 1 editor: Adam Bicz human_review: null version: null --- As of: 2026-10-09 19:24 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/e/2026-10-08-openproblembench-82-open-problems/ # OpenProblemBench: 82 unresolved math and theoretical-physics problems Full title: OpenProblemBench: 82 unresolved math and theoretical-physics problems; GPT-6 Astra has the highest model-judged solve rate (14%) On 8 Oct 2026 a CAS/USTC team (corresponding author Kun Chen) released OpenProblemBench (arXiv 2610.11118): 82 open problems from the mathematics and theoretical-physics literature, each with context and prior progress, chosen so that proposed solutions can be checked comparatively clearly. Four evaluator models (GPT-5.6 Sol, Kimi K3, GLM-5.3, Qwen3.8-Max) grade each submission without reference solutions. GPT-6 Astra (in Codex CLI) reached a mean judged solve rate of 14.0%, against 5.5–6.7% for full-size open models and 2.4–3.7% for Flash models. "Solved" here means judged solved by models, not verified by experts. ## Key facts - 82 problems; 7 solver configurations; 2,296 reviews; four evaluator models; outcome categories: solved, breakthrough partial, nontrivial partial, trivial partial, incomplete - Solvers: GPT-6 Astra and GPT-5.6 Sol in Codex CLI; DeepSeek-V4.1-Flash in DeepSeek Harness; Kimi K3, GLM-5.3, Qwen3.8-Max and GLM-5.3-Flash in OpenCode; an extra Qwen run compares OpenCode with Claude Code - GPT-6 Astra leads under every evaluator (15, 13, 9 and 9 solved); the union across configurations contains 13–16 problems judged solved by a fixed evaluator - Qwen3.8-Max and Kimi K3 get solved or substantive-partial judgments in about 73–74% of reviews, GPT-5.6 Sol 63%; DeepSeek-V4.1-Flash is dominated by trivial partial outcomes (64.6%) ## What happened OpenProblemBench evaluates models on questions with no known answer. Grading is by other models, which check the stated obligations, quantifiers and decisive steps. The authors acknowledge that this measures judged progress, not expert-verified solutions. ## Why it matters It adds a benchmark of open research problems, after FrontierMath-style tests with known answers. It ranks the AI systems that are producing real conjecture results in 2026, with GPT-6 Astra clearly ahead and Chinese open models close to each other. ## Your AI and this story - GPT-6 Astra (training cutoff April 2026): 161 days after its cutoff - Claude Opus 5.5 (training cutoff June 2026): 100 days after its cutoff - Gemini 3.8 Flash (training cutoff March 2026): 191 days after its cutoff - Grok 4.7 (training cutoff May 2026): 130 days after its cutoff ## Sources 1. [OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences (arXiv 2610.11118)](https://arxiv.org/abs/2610.11118) (arxiv.org, paper) ## Changes - 2026-10-09 (filed): Created from the arXiv PDF ## Related - 2026-10-06: [OpenAI releases 722 AI-written math manuscripts claiming hundreds of open problems](https://postcutoff.com/e/2026-10-06-openai-math-release-722-manuscripts/index.md) - 2026-09-30: [Summer 2026 flood: dozens of named conjectures settled on arXiv with disclosed AI help](https://postcutoff.com/e/2026-09-30-ai-assisted-conjecture-wave-summer-2026/index.md)