BixBench3 tests AI agents on whole computational-biology studies from raw data; best score 0.48 (GPT-5.6 Sol)
BixBench3 (arXiv 2608.25286, Aug 26, 2026), from Edison Scientific (Zane Koch, Jon M. Laurent, Samuel G. Rodriques, Andrew D. White and colleagues), gives an agent a research objective, methods guidance from a published biology paper, and that study's raw data. The agent must reproduce the paper's results. Across 20 studies and 138 graded artifacts, 13 frontier models scored between 0.00 and 0.48, and they did much worse on datasets over 100 GB.
Key facts
- 20 tasks from published studies, 17 assay types, 9 scientific domains, 138 graded artifacts compared with the originals
- Top scores: GPT-5.6 Sol 0.48, Kimi K3 0.47, GLM 5.2 0.46, Claude Opus 4.8 0.46, Claude Opus 5 0.41; lowest Gemini 3.1 Flash Lite 0.00
- Data scale hurts: ~0.36–0.37 on smaller datasets vs 0.10 on tasks over 100 GB; analyses of depth ≥3 dropped to 0.24
- Cost per attempt: on average 6.8 hours, 102M tokens and $43; maximum 24 hours, 1.07B tokens and $525. Better models used fewer resources
- Common failures: premature termination and repetitive retry loops (~2× more frequent in the worst attempts)
- Public harness and deterministic grader at github.com/EdisonScientific/BixBench3; 20-task dataset on Hugging Face; v2 posted Sept 22, 2026
- Successor to BixBench (arXiv 2503.00096, 2025), the benchmark OpenAI cited for GPT-Rosalind
What happened
BixBench3 moves biology-agent evaluation from question answering to reproducing whole studies from raw data, the work a computational biologist actually does. Frontier agents got roughly halfway on the best runs, and large data volumes and long analysis chains broke them.
Why it matters
It is an end-to-end test for "AI scientist" claims in biology. Its finding that open models from China (Kimi K3, GLM 5.2) score close to the top closed model is also notable.
Changelog
- 2026-10-04: created (resolves leads.md line)
Related events
- OpenAI releases LifeSciBench, 750 expert-written life-science research tasks; GPT-Rosalind leads with a 36% pass rate ★★
- OpenAI launches GPT-Rosalind, a trusted-access reasoning model for life-sciences research ★★★
- Edison Scientific's Kosmos AI scientist claims six months of research per run ★★★
Sources (5)
- paperarXiv 2608.25286: BixBench3, benchmarking AI agents on research-study-scale computational biology tasks
- codeGitHub: EdisonScientific/BixBench3 (harness and grader)
- officialEdison Advances: BixBench3 benchmark page
- paperBixBench (original), arXiv 2503.00096
- pressKiin Bio newsletter: Google's GlucoFM, Foldseek-Interface and BixBench3
id: 2026-08-26-bixbench3-computational-biology-agents · updated 2026-10-04 · open in the interactive timeline