MLPerf Inference v6.1: record 30 submitters, first peer-reviewed Vera Rubin NVL72 results, new RAG and edge-agentic tests
On Sept 16, 2026 MLCommons published MLPerf Inference v6.1 results with a record 30 submitting organizations, two new tests (End-to-End RAG and Edge Agentic Inference) and the first peer-reviewed numbers for NVIDIA's Vera Rubin NVL72 (preview), which NVIDIA says delivers up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. MLCommons also said the new MLPerf Endpoints suite will replace Inference for the datacenter.
Key facts
- Record 30 submitting organizations, six first-time submitters (Atlas Inference, Crusoe, Orrick, ScitiX, VibeHPC, Naeem Khoshnevis)
- New tests: End-to-End RAG (embedding → retriever → re-ranker → LLMs) and Edge Agentic Inference (multi-turn agentic workloads at the edge; NVIDIA ran it with Qwen3.6-27B on Jetson AGX Thor)
- Speculative decoding now allowed in the interactive scenario for two benchmarks, plus GPT-OSS
- Five new processors/accelerators: AMD Ryzen AI Max+ 395, AMD Instinct MI350P, Intel Arc Pro B70 (available); NVIDIA Rubin and Vera Rubin NVL72 (preview)
- Largest system ever submitted to MLPerf Inference: 512 accelerators; plus a heterogeneous two-vendor system and one geographically distributed across the Pacific Ocean
- Best per-accelerator server result: VLM test up 2.99x vs v6.0 (six months earlier); DeepSeek-R1 test up 5.7x vs v5.1 (one year earlier)
- NVIDIA: Vera Rubin NVL72 up to 3.7x GB300 NVL72 throughput on Qwen3-VL (vLLM + Dynamo) and up to 2.5x on DeepSeek-R1 (TensorRT-LLM); DeepSeek-R1 scaled from 1 to 4 GB300 NVL72 racks (288 GPUs) at 99% scaling efficiency
- Over 50% of submitters used MLPerf's new API-centric harness; MLCommons says MLPerf Endpoints 'will replace Inference in our family of benchmarks for the datacenter'
What happened
MLCommons released MLPerf Inference v6.1 on Sept 16, 2026. The round set a participation record (30 organizations) and added two tests aimed at multi-step deployments: an End-to-End RAG pipeline and an Edge Agentic Inference test for multi-turn agentic workloads on single-user edge devices. Speculative decoding became an allowed optimization in some interactive scenarios.
NVIDIA submitted preview results for Vera Rubin NVL72 on DeepSeek-R1 and Qwen3-VL, the platform's first peer-reviewed benchmark numbers, claiming up to 3.7x (Qwen3-VL) and 2.5x (DeepSeek-R1) the throughput of GB300 NVL72. Nebius also submitted Vera Rubin NVL72 preview results. AMD (Instinct MI350P, Ryzen AI Max+ 395) and Intel (Arc Pro B70) debuted new chips. NVIDIA additionally cited a 30x gain over GB300 NVL72 in SemiAnalysis's AgentX benchmark in preview testing (not an MLPerf result).
MLCommons said its new API-centric harness, used by over half of submitters, is the basis of "MLPerf Endpoints", which will replace MLPerf Inference for datacenter systems.
Why it matters
MLPerf is the main audited cross-vendor inference benchmark. This round gave the first independent check on Vera Rubin's claimed generational gains, and its new tests show the benchmark following the industry from single-shot chat toward RAG pipelines and agents.
Changelog
- 2026-10-02: created (found via NVIDIA blog while checking lab news).
Related events
- NVIDIA posts $96.2B quarter; Vera Rubin in full production and deploying at major clouds ★★★★
- NVIDIA GTC 2026: Vera Rubin platform, Groq 3 LPX, Feynman preview and $1T demand outlook ★★★★
Sources (7)
- officialMLCommons: Participation record with new MLPerf Inference v6.1 results
- pressGlobeNewswire: MLCommons press release (Sept 16, 2026)
- officialMLCommons: Where the industry is investing — a look at MLPerf Inference v6.1
- officialNVIDIA: Vera Rubin NVL72 delivers leading performance in MLPerf Inference v6.1 debut
- officialNebius: MLPerf Inference v6.1 results on Vera Rubin NVL72
- officialAMD ROCm blog: MLPerf Inference v6.1 submission
- pressTechTimes: MLPerf v6.1 puts first peer-reviewed numbers on Vera Rubin
id: 2026-09-16-mlperf-inference-v6-1 · updated 2026-10-02 · open in the interactive timeline