Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. BixBench3 tests AI agents on whole computational-biology…

BixBench3 tests AI agents on whole computational-biology studies from raw data; best score 0.48 (GPT-5.6 Sol)

★★after cutoffbenchmarkEdison Scientificconfidence: high

BixBench3 (arXiv 2608.25286, Aug 26, 2026), from Edison Scientific (Zane Koch, Jon M. Laurent, Samuel G. Rodriques, Andrew D. White and colleagues), gives an agent a research objective, methods guidance from a published biology paper, and that study's raw data. The agent must reproduce the paper's results. Across 20 studies and 138 graded artifacts, 13 frontier models scored between 0.00 and 0.48, and they did much worse on datasets over 100 GB.

Key facts

What happened

BixBench3 moves biology-agent evaluation from question answering to reproducing whole studies from raw data, the work a computational biologist actually does. Frontier agents got roughly halfway on the best runs, and large data volumes and long analysis chains broke them.

Why it matters

It is an end-to-end test for "AI scientist" claims in biology. Its finding that open models from China (Kimi K3, GLM 5.2) score close to the top closed model is also notable.

Changelog

  • 2026-10-04: created (resolves leads.md line)

Related events

  1. OpenAI releases LifeSciBench, 750 expert-written life-science research tasks; GPT-Rosalind leads with a 36% pass rate ★★
  2. OpenAI launches GPT-Rosalind, a trusted-access reasoning model for life-sciences research ★★★
  3. Edison Scientific's Kosmos AI scientist claims six months of research per run ★★★

Sources (5)

id: 2026-08-26-bixbench3-computational-biology-agents · updated 2026-10-04 · open in the interactive timeline