Arena raises $200M at $3.1B and launches an AI Alignment Index for agents
OpenAI models lead, Claude Opus 5.5 is 6th
Confirmed
The takeaway
On Oct 8, 2026 Arena, the crowdsourced model-leaderboard company that began as UC Berkeley’s Chatbot Arena, raised a $200M Series B at a $3.1B valuation (up from $1.7B in January), led by Lightspeed and Khosla.
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 3 of 5
- Last verified
- 8 October 2026
Your AI and this story
- GPT-6 Astra161 days after its cutoff
- Claude Opus 5.5100 days after its cutoff
- Gemini 3.8 Flash191 days after its cutoff
- Grok 4.7130 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 100 days before it.
Key facts
- Series B: $200M at $3.1B, led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z and Felicis (TechCrunch)
- Previous round: $150M Series A at $1.7B (January 2026); annualized revenue ~$30M in January and $100M by June 2026 (TechCrunch)
- Alignment Index measures three failure types in real Agent Arena sessions: unauthorized actions, false attribution (misrepresenting what the user said) and deceptive completion (reporting unfinished tasks as done)
- First edition (dated Sept 30, 2026; 72,509 sessions; 27 models): 1 GPT-6.1 Sol 87.9, 2 GPT-6 Astra 87.8, 3 GPT-6 Luna 87.8, 4 GPT-6 Sol 87.6, 5 GPT-5.6 Sol 84.2, 6 Claude Opus 5.5 83.2, 7 Grok 4.7 82.7, 8 Grok 4.6 80.8, 9 Claude Fable 5.1 80.1, 10 Claude Sonnet 5.5 High 79.5 (arena.ai leaderboard, ±1–2 points)
- Arena’s pitch: ‘static benchmarks break down once models recognize they’re being tested’ (TechCrunch)
What happened
Arena announced its Series B together with a new agent-behaviour leaderboard. Unlike its preference-vote rankings, the Alignment Index is computed from failures observed in real user sessions in Agent Arena. The leaderboard page gives scores with confidence intervals; the blog post with the full method was not readable by our fetcher.
Why it matters
Arena’s rankings are widely quoted in model launches. An index of agent honesty and permission-following from live sessions is a new kind of public alignment signal, and its first edition ranks OpenAI’s GPT-6 models above Anthropic’s.
Sources
4 sources from 3 sites. Numbers match the chips in the text.
4 sources: 2 primary, 2 press
Primary
- Arena: AI Alignment Index leaderboard (agents)arena.ai, official
- Arena blog: AI Alignment Indexarena.ai, official
Press
- TechCrunch: Popular AI leaderboard Arena nearly doubles valuation to $3.1B in 10 monthstechcrunch.com, press
- Bloomberg: AI model evaluator Arena valued at over $3 billion in roundbloomberg.com, press
Changes
- Filed