Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals
On Sept 28, 2026 Artificial Analysis launched the Cyber Index, which combines three benchmarks of how well AI agents find and patch vulnerabilities, and the Cyber Index Alliance with Collinear AI, IBM, NVIDIA and Vercel to set shared evaluation standards for defensive cyber tasks. The best models solved only 41% of expert-verified vulnerabilities on its DeepsecBench-AA component.
Key facts
- Components: CWE-Bench-AA (120 tasks: audit open-source repos and patch OWASP Top 10 flaws), DeepsecBench-AA (find vulnerabilities vs expert-verified findings), CyberGym-E2E-AA (131 tasks: discover, reproduce and patch C/C++ memory-safety bugs)
- All runs use Stirrup, an open-source agent harness
- Best models solve 41% of expert-verified vulnerabilities in DeepsecBench-AA; '55% of failed attempts fixed the primary issue but left a related one open'
- Refusal rates above 98% for frontier models on CyberGym-E2E-AA tasks
- Launch partners: Collinear AI, IBM, NVIDIA, Vercel
What happened
Artificial Analysis, known for its model leaderboards, added a defensive-cyber index and recruited vendors to agree on how such work is measured.
Why it matters
Cyber capability is now the main axis of frontier-model risk debates. A shared, defense-oriented benchmark from an independent evaluator gives buyers and policymakers a common yardstick. The high refusal rates show that safety filters strongly shape these scores.
Changelog
- 2026-09-30: created (sweep 2026-09-29)
Related events
Sources (1)
id: 2026-09-28-artificial-analysis-cyber-index · updated 2026-09-30 · open in the interactive timeline