Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. MLPerf Inference v6.1: record 30 submitters, first…

MLPerf Inference v6.1: record 30 submitters, first peer-reviewed Vera Rubin NVL72 results, new RAG and edge-agentic tests

★★★after cutoffbenchmarkMLCommonsNVIDIAconfidence: high

On Sept 16, 2026 MLCommons published MLPerf Inference v6.1 results with a record 30 submitting organizations, two new tests (End-to-End RAG and Edge Agentic Inference) and the first peer-reviewed numbers for NVIDIA's Vera Rubin NVL72 (preview), which NVIDIA says delivers up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. MLCommons also said the new MLPerf Endpoints suite will replace Inference for the datacenter.

Key facts

What happened

MLCommons released MLPerf Inference v6.1 on Sept 16, 2026. The round set a participation record (30 organizations) and added two tests aimed at multi-step deployments: an End-to-End RAG pipeline and an Edge Agentic Inference test for multi-turn agentic workloads on single-user edge devices. Speculative decoding became an allowed optimization in some interactive scenarios.

NVIDIA submitted preview results for Vera Rubin NVL72 on DeepSeek-R1 and Qwen3-VL, the platform's first peer-reviewed benchmark numbers, claiming up to 3.7x (Qwen3-VL) and 2.5x (DeepSeek-R1) the throughput of GB300 NVL72. Nebius also submitted Vera Rubin NVL72 preview results. AMD (Instinct MI350P, Ryzen AI Max+ 395) and Intel (Arc Pro B70) debuted new chips. NVIDIA additionally cited a 30x gain over GB300 NVL72 in SemiAnalysis's AgentX benchmark in preview testing (not an MLPerf result).

MLCommons said its new API-centric harness, used by over half of submitters, is the basis of "MLPerf Endpoints", which will replace MLPerf Inference for datacenter systems.

Why it matters

MLPerf is the main audited cross-vendor inference benchmark. This round gave the first independent check on Vera Rubin's claimed generational gains, and its new tests show the benchmark following the industry from single-shot chat toward RAG pipelines and agents.

Changelog

  • 2026-10-02: created (found via NVIDIA blog while checking lab news).

Related events

  1. NVIDIA posts $96.2B quarter; Vera Rubin in full production and deploying at major clouds ★★★★
  2. NVIDIA GTC 2026: Vera Rubin platform, Groq 3 LPX, Feynman preview and $1T demand outlook ★★★★

Sources (7)

id: 2026-09-16-mlperf-inference-v6-1 · updated 2026-10-02 · open in the interactive timeline