Toby Ord: 'Swarm Scaling' — how agent-swarm performance scales with swarm size (λ ≈ 0.48–0.68)
Toby Ord @tobyordoxford · blog · 2026-09-21 · ★★★ · archived
First quantitative estimate of the 'parallelizability' of AI agent swarms, from OpenAI's GPT-5.6 swarm charts; Ord argues the measured λ keeps recursive self-improvement scenarios plausible. Featured in Import AI 475.
Summary
Oxford philosopher Toby Ord analyses the multi-agent swarm charts in OpenAI's GPT-5.6 launch post (BrowseComp, SEC-Bench Pro, Terminal-Bench). He fits a "stepping on toes" parallelization parameter λ, borrowed from the economics of research teams. With N agents, a task finishes about N^λ times faster but uses about N^(1−λ) times more compute. His estimates: BrowseComp 0.68, SEC-Bench Pro 0.57, Terminal-Bench 0.48. In practice, a 4x larger swarm needs about 2x the total tokens for the same score. Putting extra inference into swarm size gives about half the score gain of putting it into longer chain of thought, so the scale-up must be "squared". Swarms pay off mainly in speed. Ord argues that high λ makes intelligence explosions more likely, and that the measured values match the ranges assumed in earlier RSI models. He urges tracking λ as orchestration improves. The X launch thread (Sept 21, ~74k views) was later discussed by Peter J. Liu. Jack Clark's Import AI 475 (Oct 5) led with the essay.
Archived text
(blog: summary plus short quotes only)
"You need to do the scale-up of compute twice to get to the same capability, squaring the total multiplier needed." "The most important answer is speed." "High λ makes intelligence explosions more likely."
Thread: https://x.com/tobyordoxford/status/2102111263660642673 · Import AI 475: https://jack-clark.net/2026/10/05/import-ai-475-swarm-scaling-google-deepmind-watermarks-biology-and-the-ai-science-economy/
People
Related events
All posts · id: 2026-09-21-toby-ord-swarm-scaling