OpenAI releases LifeSciBench, 750 expert-written life-science research tasks; GPT-Rosalind leads with a 36% pass rate
On June 17, 2026 OpenAI introduced LifeSciBench, a benchmark of 750 expert-authored tasks that span seven research workflows and seven biological domains, graded with 19,020 rubric criteria. Its life-sciences model GPT-Rosalind scored best but passed only 36.1% of tasks, and 22.8% of tasks were passed by no model, so the benchmark is far from saturated.
Key facts
- 750 tasks written by 173 PhD-level scientists and validated by 453 reviewers (97% with doctorates); ~25 rubric criteria per task, 19,020 in total
- About 79% of tasks need multiple reasoning steps (four on average); ~53% include attached artifacts such as figures, PDFs or sequence files
- Task pass rates (single-turn, internet access): GPT-Rosalind 36.1%, GPT-5.5 25.7%, Gemini 3.1 Pro 23.6%, GPT-5.4 20.7%, Grok 4.3 13.0%
- Normalized scores: GPT-Rosalind 0.576, GPT-5.5 0.519, Gemini 3.1 Pro 0.515, GPT-5.4 0.479, Grok 4.3 0.399
- 171 tasks (22.8%) were passed by no model; 261 tasks had a best-model pass rate below 20%
- Artifacts are the main bottleneck: GPT-Rosalind fell from 45.1% on text-only tasks to 28.1% on tasks with attached files
- OpenAI also reported an in-house LabWorkBench result (wet-lab protocol support): GPT-Rosalind 63.2% vs GPT-5.5 55.8% (via Labcritics)
- Unverified: whether the task set is publicly released (not stated in the sources read); OpenAI's announcement page could not be fetched (403)
What happened
OpenAI released LifeSciBench, with a preprint, as a benchmark of realistic life-science research work rather than exam questions. It was published two months after GPT-Rosalind, which led the results.
Why it matters
It is a large, expert-graded measure of AI for biology research. Its main finding is that models still struggle to read real scientific data files, the skill research agents need most. As OpenAI built the benchmark and its own model leads it, independent replication matters.
Changelog
- 2026-10-04: created (resolves leads.md line)
Related events
- OpenAI launches GPT-Rosalind, a trusted-access reasoning model for life-sciences research ★★★
- BixBench3 tests AI agents on whole computational-biology studies from raw data; best score 0.48 (GPT-5.6 Sol) ★★
Sources (4)
- paperOpenAI: LifeSciBench preprint (PDF)
- pressLabcritics: LifeSciBench, OpenAI's hard new life-science benchmark, and how GPT-Rosalind stacks up (Jun 19, 2026)
- pressAI TL;DR: LifeSciBench, OpenAI's 750-task benchmark
- pressAI Weekly: OpenAI's LifeSciBench tests AI on 750 life-science research tasks
id: 2026-06-17-openai-lifescibench · updated 2026-10-04 · open in the interactive timeline