As of: 2026-10-10 14:45 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/news/benchmark/ # AI news: Benchmarks 24 events in Benchmarks of 1,105 in the log, newest first. Page 1 of 1, 50 events per page, grouped by the day each event happened. ## Thursday 8 October 2026 - [Arena raises $200M at $3.1B and launches an AI Alignment Index for agents](https://postcutoff.com/e/2026-10-08-arena-200m-3-1b-alignment-index/) (Benchmarks; Arena). On Oct 8, 2026 Arena, the crowdsourced model-leaderboard company that began as UC Berkeley's Chatbot Arena, raised a $200M Series B at a $3.1B valuation (up from $1.7B in January), led by Lightspeed and Khosla. Source: https://arena.ai/leaderboard/agent/alignment - [OpenProblemBench: 82 unresolved math and theoretical-physics problems](https://postcutoff.com/e/2026-10-08-openproblembench-82-open-problems/) (Benchmarks; Institute of Theoretical Physics CAS, University of Science and Technology of China). It adds a benchmark of open research problems, after FrontierMath-style tests with known answers. Source: https://arxiv.org/abs/2610.11118 ## Thursday 1 October 2026 - [Open-source ARC-AGI-3 agents top out at 27.9% at ARC Prize 2026 Milestone #2](https://postcutoff.com/e/2026-10-01-arc-prize-2026-milestone-2-winners/) (Benchmarks; ARC Prize Foundation, Tufa Labs). The Kaggle track runs open models under Kaggle's compute limits, so it measures how far open methods get without frontier APIs. Source: https://x.com/arcprize/status/2105737436450201734 ## Monday 28 September 2026 - [Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals](https://postcutoff.com/e/2026-09-28-artificial-analysis-cyber-index/) (Benchmarks; Artificial Analysis, IBM, NVIDIA, Vercel, Collinear AI). Cyber capability is now the main axis of frontier-model risk debates. Source: https://artificialanalysis.ai/articles/artificial-analysis-cyber-index ## Thursday 24 September 2026 - [C5R opens Facility-0, an AI-run wet lab, and the SciUniverse benchmark](https://postcutoff.com/e/2026-09-24-c5r-facility-0-sciuniverse-benchmark/) (Benchmarks; C5R). Most AI-for-science benchmarks test reasoning on text or data. Source: https://c5r.net/sciuniverse/ ## Wednesday 23 September 2026 - [OpenAI releases MentalHealthBench, an open benchmark for AI in mental-health conversations](https://postcutoff.com/e/2026-09-23-openai-mentalhealthbench/) (Benchmarks; OpenAI). Mental-health use of chatbots was a major 2025–2026 safety and litigation topic, and this gives labs and regulators a shared measure. Source: https://openai.com/index/introducing-mentalhealthbench/ - [DrivingBench: GPT-6 Astra is the first frontier LLM to complete a cone course driving a real Toyota Corolla](https://postcutoff.com/e/2026-09-23-drivingbench-llms-drive-real-car/) (Benchmarks; OpenAI). A small but vivid probe of general models' embodied, closed-loop control. Source: https://drivingbench.com/ ## Tuesday 22 September 2026 - [CAIS releases HLE-Diamond, a refined 1,000-question Humanity's Last Exam](https://postcutoff.com/e/2026-09-22-hle-diamond/) (Benchmarks; Center for AI Safety, Scale AI). If confirmed, the benchmark once billed as 'the last exam' is close to saturation within about 20 months of release. Source: https://lastexam.ai/blog/hle-diamond ## Wednesday 16 September 2026 - [MLPerf Inference v6.1 draws a record 30 submitters and the first Vera Rubin NVL72 results](https://postcutoff.com/e/2026-09-16-mlperf-inference-v6-1/) (Benchmarks; MLCommons, NVIDIA). MLPerf is the main audited cross-vendor inference benchmark. Source: https://mlcommons.org/2026/09/mlperf-inference-v6-1-results/ ## Thursday 3 September 2026 - [GPT-6 Astra scores 62.7% on ARC-AGI-3, outacting humans on 96% of levels](https://postcutoff.com/e/2026-09-03-arc-agi-3-gpt-6-astra/) (Benchmarks; ARC Prize Foundation, OpenAI; historic). ARC-AGI-3 was meant to measure human-like skill acquisition; its near-saturation (and the harness gap) shows both how fast agentic reasoning improved in 2026 and how much scaffolding now drives scores. Source: https://arcprize.org/blog/astra - [Qwen and Taobao release E-Commerce Bench](https://postcutoff.com/e/2026-09-03-qwen-e-commerce-bench/) (Benchmarks; Alibaba, Qwen, Taobao & Tmall Group). Results differed by orders of magnitude, and the top earner was not the most careful operator. Source: https://qwen.ai/blog?id=e-commerce-bench ## Wednesday 26 August 2026 - [BixBench3 tests AI agents on whole computational-biology studies from raw data](https://postcutoff.com/e/2026-08-26-bixbench3-computational-biology-agents/) (Benchmarks; Edison Scientific). It is an end-to-end test for "AI scientist" claims in biology. Source: https://arxiv.org/abs/2608.25286 ## Monday 24 August 2026 - [Artificial Analysis launches the Speech Agent Arena for speech-to-speech voice agents](https://postcutoff.com/e/2026-08-24-artificial-analysis-speech-agent-arena/) (Benchmarks; Artificial Analysis). Voice agents are being sold into customer service, where finishing the task matters more than sounding natural. Source: https://artificialanalysis.ai/articles/announcing-the-speech-agent-arena ## Friday 24 July 2026 - [Project Pilot: Anthropic and Andon Labs test whether AI models can fly a surveillance drone (Drone-Bench)](https://postcutoff.com/e/2026-07-24-anthropic-andon-project-pilot-drone-bench/) (Benchmarks; Anthropic, Andon Labs). Autonomous aerial surveillance is a clearly dual-use capability, and the study suggests frontier models were close to it by mid-2026, except for 3D mapping. Source: https://www.anthropic.com/research/project-pilot ## Wednesday 17 June 2026 - [OpenAI releases LifeSciBench, 750 expert-written life-science research tasks](https://postcutoff.com/e/2026-06-17-openai-lifescibench/) (Benchmarks; OpenAI). It is a large, expert-graded measure of AI for biology research. Source: https://cdn.openai.com/pdf/b4299379-0a97-4ffa-8b9b-c3fbb299caa9/lifescibench_preprint.pdf ## Wednesday 10 June 2026 - [AI systems pass 7 of 10 unpublished research problems in First Proof's second batch](https://postcutoff.com/e/2026-06-10-first-proof-second-batch/) (Benchmarks; First Proof Foundation, ETH Zurich, OpenAI, UCLA, Princeton University). It is the most carefully refereed measurement so far of AI on genuine research problems. Source: https://1stproof.org/assets/docs/report.pdf ## Wednesday 25 March 2026 - [ARC Prize launches ARC-AGI-3, an interactive game benchmark where frontier AI scored under 1%](https://postcutoff.com/e/2026-03-25-arc-agi-3-launch/) (Benchmarks; ARC Prize Foundation; major). It was designed as the hardest-to-game AGI benchmark of 2026; within six months it was largely cracked (see GPT-6 Astra entry), illustrating the pace of agentic progress. Source: https://arcprize.org/arc-agi/3 ## Wednesday 11 February 2026 - [Ai2 launches MolmoSpaces, an open simulation ecosystem and leaderboard for generalist robot policies](https://postcutoff.com/e/2026-02-11-ai2-molmospaces/) (Benchmarks; Ai2). Robot learning lacked shared benchmarks like those language models have. Source: https://allenai.org/blog/molmospaces ## Thursday 29 January 2026 - [METR releases Time Horizon 1.1 with expanded long-task suite](https://postcutoff.com/e/2026-01-29-metr-time-horizon-1-1/) (Benchmarks; METR). As frontier horizons approach the top of the suite, METR's own caveat (unreliable >16h) signals the benchmark itself is near saturation. Source: https://metr.org/blog/2026-1-29-time-horizon-1-1/ ## Saturday 6 December 2025 - [Arc Institute announces first Virtual Cell Challenge winners](https://postcutoff.com/e/2025-12-06-virtual-cell-challenge-2025-winners/) (Benchmarks; Arc Institute, BioMap, Altos Labs, NVIDIA). Virtual cells are a major goal of AI biology, and this challenge is becoming the field's shared benchmark. Source: https://arcinstitute.org/news/virtual-cell-challenge-2025-wrap-up ## Wednesday 17 September 2025 - [AI reaches gold-medal level at the ICPC World Finals](https://postcutoff.com/e/2025-09-17-icpc-gold-ai/) (Benchmarks; OpenAI, Google DeepMind; major). Following IMO gold, confirmed elite-human-level algorithmic problem solving by general-purpose reasoning models. Source: https://deepmind.google/blog/gemini-achieves-gold-medal-level-at-the-international-collegiate-programming-contest-world-finals/ ## Sunday 22 June 2025 - [RoboArena: crowd-sourced, double-blind real-world evaluation of generalist robot policies](https://postcutoff.com/e/2025-06-22-roboarena/) (Benchmarks; RoboArena consortium). Real-world robot evaluation is expensive and hard to standardize. Source: https://arxiv.org/abs/2506.18123 ## Friday 20 December 2024 - [OpenAI announces o3, scoring 75.7–87.5% on ARC-AGI](https://postcutoff.com/e/2024-12-20-openai-o3/) (Benchmarks; OpenAI, ARC Prize; historic). Convinced many observers that reasoning models were on a steep trajectory; ARC Prize called it a genuine step-change. Source: https://arcprize.org/blog/oai-o3-pub-breakthrough ## June 2009, day not recorded - [ImageNet dataset presented at CVPR 2009](https://postcutoff.com/e/2009-06-20-imagenet/) (Benchmarks; Princeton University, Stanford University; historic). Showed that data scale was a key ingredient of progress; AlexNet's 2012 ImageNet win kicked off the deep learning era. Source: https://doi.org/10.1109/CVPR.2009.5206848 Other views: All https://postcutoff.com/news/; Major only https://postcutoff.com/news/major/; Policy & safety https://postcutoff.com/news/policy-safety/; Science & math https://postcutoff.com/news/science/; Business https://postcutoff.com/news/business/; Model releases https://postcutoff.com/news/model-release/; Research https://postcutoff.com/news/research/; Chips & compute https://postcutoff.com/news/hardware-compute/; Products https://postcutoff.com/news/product/; Open source https://postcutoff.com/news/open-source/; Agents https://postcutoff.com/news/agents/; Robotics https://postcutoff.com/news/robotics/; Media generation https://postcutoff.com/news/media-generation/; Benchmarks https://postcutoff.com/news/benchmark/; Culture https://postcutoff.com/news/culture/; Milestones https://postcutoff.com/news/milestone/. Feeds: https://postcutoff.com/feeds/benchmark.xml