{"schema":"postcutoff/event@1","as_of":"2026-10-09T19:24:00+02:00","url":"https://postcutoff.com/e/2026-10-08-openproblembench-82-open-problems/","md":"https://postcutoff.com/e/2026-10-08-openproblembench-82-open-problems/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"2026-10-08-openproblembench-82-open-problems","date":"2026-10-08","date_precision":"day","short_title":"OpenProblemBench: 82 unresolved math and theoretical-physics problems","deck":"GPT-6 Astra has the highest model-judged solve rate (14%)","takeaway":null,"category":"benchmark","category_label":"Benchmarks","importance":2,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":[{"n":1,"title":"OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences (arXiv 2610.11118)","url":"https://arxiv.org/abs/2610.11118","type":"paper","group":"primary","domain":"arxiv.org"}],"official":1,"filed":"2026-10-09","updated":"2026-10-09","orgs":["Institute of Theoretical Physics CAS","University of Science and Technology of China"],"title":"OpenProblemBench: 82 unresolved math and theoretical-physics problems; GPT-6 Astra has the highest model-judged solve rate (14%)","summary":"On 8 Oct 2026 a CAS/USTC team (corresponding author Kun Chen) released OpenProblemBench (arXiv 2610.11118): 82 open problems from the mathematics and theoretical-physics literature, each with context and prior progress, chosen so that proposed solutions can be checked comparatively clearly. Four evaluator models (GPT-5.6 Sol, Kimi K3, GLM-5.3, Qwen3.8-Max) grade each submission without reference solutions. GPT-6 Astra (in Codex CLI) reached a mean judged solve rate of 14.0%, against 5.5–6.7% for full-size open models and 2.4–3.7% for Flash models. \"Solved\" here means judged solved by models, not verified by experts.","key_facts":["82 problems; 7 solver configurations; 2,296 reviews; four evaluator models; outcome categories: solved, breakthrough partial, nontrivial partial, trivial partial, incomplete","Solvers: GPT-6 Astra and GPT-5.6 Sol in Codex CLI; DeepSeek-V4.1-Flash in DeepSeek Harness; Kimi K3, GLM-5.3, Qwen3.8-Max and GLM-5.3-Flash in OpenCode; an extra Qwen run compares OpenCode with Claude Code","GPT-6 Astra leads under every evaluator (15, 13, 9 and 9 solved); the union across configurations contains 13–16 problems judged solved by a fixed evaluator","Qwen3.8-Max and Kimi K3 get solved or substantive-partial judgments in about 73–74% of reviews, GPT-5.6 Sol 63%; DeepSeek-V4.1-Flash is dominated by trivial partial outcomes (64.6%)"],"key_numbers":[],"tags":["benchmark","math","physics","open-problems","gpt-6-astra","llm-as-judge","open-weights"],"science":null,"body_md":"## What happened\n\nOpenProblemBench evaluates models on questions with no known answer. Grading is by other models, which check the stated\nobligations, quantifiers and decisive steps. The authors acknowledge that this measures judged progress, not expert-verified\nsolutions.\n\n## Why it matters\n\nIt adds a benchmark of open research problems, after FrontierMath-style tests with known answers. It ranks the AI systems\nthat are producing real conjecture results in 2026, with GPT-6 Astra clearly ahead and Chinese open models close to each\nother.","disputed":[],"related":[{"id":"2026-10-06-openai-math-release-722-manuscripts","url":"https://postcutoff.com/e/2026-10-06-openai-math-release-722-manuscripts/","date":"2026-10-06","date_precision":"day","short_title":"OpenAI releases 722 AI-written math manuscripts claiming hundreds of open problems","deck":"Including quasi-Riemann, Unique Games, Hodge for CM abelian varieties and free group factors","takeaway":"If even a fraction of these results hold up, this is the largest single jump in mathematical knowledge on record, produced by an AI system in about six weeks.","category":"science","category_label":"Science & math","importance":5,"confidence":"high","status":{"key":"pending","labels":["Event confirmed","Awaiting review"]},"sources":58,"official":12,"filed":"2026-10-07","updated":"2026-10-09","orgs":["OpenAI"]},{"id":"2026-09-30-ai-assisted-conjecture-wave-summer-2026","url":"https://postcutoff.com/e/2026-09-30-ai-assisted-conjecture-wave-summer-2026/","date":"2026-09-30","date_precision":"day","short_title":"Summer 2026 flood: dozens of named conjectures settled on arXiv with disclosed AI help","deck":null,"takeaway":"This entry catalogues about 50 of them, with the AI role as the authors state it.","category":"science","category_label":"Science & math","importance":4,"confidence":"medium","status":{"key":"pending","labels":["Awaiting review"]},"sources":9,"official":4,"filed":"2026-09-30","updated":"2026-10-09","orgs":["OpenAI","Anthropic","Google DeepMind","various mathematicians"]}],"people":[],"posts":[],"videos":[],"models":[],"changes":[{"date":"2026-10-09","type":"filed","text":"Created from the arXiv PDF"}],"provenance":{"agents":[{"model":"Claude Opus 5.5","maker":"Anthropic","tool":"Claude Code"}],"filed":"2026-10-09","run":null,"sources_read":"The arXiv PDF","updated":"2026-10-09","human_review":null,"version":null},"gaps":[{"model_id":"gpt-6-astra","name":"GPT-6 Astra","cutoff":"2026-04","days_after":161,"in_training_data":false},{"model_id":"claude-opus-5-5","name":"Claude Opus 5.5","cutoff":"2026-06","days_after":100,"in_training_data":false},{"model_id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","cutoff":"2026-03","days_after":191,"in_training_data":false},{"model_id":"grok-4-7","name":"Grok 4.7","cutoff":"2026-05","days_after":130,"in_training_data":false}],"short_url":null}