NVIDIA Vera Rubin NVL72 Posts Up to 3.7x Throughput Gain
NVIDIA's Vera Rubin NVL72 made its first MLPerf Inference preview submission with results that the company says outperform its GB300 NVL72 system by up to 3.7x. The results highlight gains from new hardware, rack-scale networking and continuing software optimization.

Vera Rubin enters MLPerf Inference
NVIDIA's Vera Rubin NVL72 system has made its debut in the MLPerf Inference v6.1 benchmark with preview results on DeepSeek-R1 and Qwen3-VL, two demanding workloads in the latest test suite. According to NVIDIA, the platform achieved as much as 3.7 times the throughput of GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios.
The Qwen3-VL tests used vLLM together with NVIDIA Dynamo, the company's open-source inference framework. On DeepSeek-R1, NVIDIA reported throughput gains of up to 2.5 times compared with GB300 NVL72 when using the TensorRT-LLM software library. Because these are preview submissions, the results represent an early view of the platform rather than a complete production benchmark record.
Hardware and software designed together
NVIDIA attributed the performance gains to full-stack design across the Vera Rubin hardware and its inference software. The system's enhanced Tensor Cores and Transformer Engine are intended to accelerate both prefill and decode, the two main stages involved in serving generative AI models.
The company also pointed to NVFP4 precision, which reduces the memory required for model weights, attention and the key-value cache. NVIDIA says the approach increases throughput while causing minimal loss of output quality. The submitted configurations made extensive use of disaggregated serving, separating prefill from decode, and large-scale expert parallelism for mixture-of-experts models such as DeepSeek-R1 and Qwen3-VL.
At the rack level, these methods rely on the Vera Rubin NVL72 scale-up domain. NVIDIA said sixth-generation NVLink and NVLink Switch provide packet rates 10 times higher and latency three times lower than off-the-shelf Ethernet, creating the interconnect foundation for the system's rack-scale operation.
GB300 demonstrates multi-rack scaling
The same MLPerf round also showed NVIDIA's GB300 NVL72 scaling from one rack to four. A DeepSeek-R1 offline submission expanded from 72 GPUs in a single rack to 288 GPUs across four racks, reaching 99% scaling efficiency, according to NVIDIA.
Scaling efficiency measures how closely throughput follows the amount of hardware added. A result near 100% indicates that the additional racks produced almost proportional gains rather than being limited by communication, orchestration or other system overhead. NVIDIA said the result depended on a combination of high-bandwidth, low-latency connections within each rack, networking between racks and request orchestration across nodes.
GB300 NVL72 also recorded results on the WAN 2.2 text-to-video benchmark. NVIDIA reported throughput of 0.65 720p videos per second at 5.7 seconds per video, describing this as nine times higher throughput and 7.5 times lower latency than a single node.
Software continues to lift results
NVIDIA said its MLPerf Inference v6.1 submissions benefited from software work completed since the previous benchmark release. GB300 NVL72 performance on Qwen3-VL improved by as much as 1.6 times compared with the company's v6.0 results.
The reported gains came from several changes, including lower key-value-cache precision, additional kernel fusion, improved kernels and disaggregated serving based on vLLM and NVIDIA Dynamo. NVIDIA added that optimization continued after the v6.1 submission deadline. Further results on GPT-OSS-120B and DLRMv3 have shown additional gains, although those figures had not yet been verified by MLCommons at the time of the announcement.
Results extend to agentic and edge workloads
NVIDIA also highlighted performance in workloads designed around AI agents, which reason, plan and act through multiple steps. In preview testing with the SemiAnalysis AgentX benchmark, the company said Vera Rubin NVL72 delivered 30 times the performance of GB300 NVL72. NVIDIA described the upcoming MLPerf Endpoints benchmark as a future way to standardize measurements for agentic inference workloads that are not fully represented by conventional throughput tests.
The benchmark participation extended beyond data-center systems. NVIDIA submitted Jetson AGX Thor results using TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B. The company also said Nebius submitted Vera Rubin NVL72 preview results and demonstrated strong performance.
NVIDIA reported participation from 19 partners in the MLPerf round, including eight using multi-node Blackwell NVL72 systems. The listed companies included cloud providers, system manufacturers and technology partners such as ASUS, Azure, Cisco, CoreWeave, Dell Technologies, Google, HPE, Oracle Cloud Infrastructure, Red Hat, Supermicro and Wiwynn. NVIDIA presented the results as evidence that performance, scaling efficiency and software development speed are central factors in the economics of AI inference infrastructure.
Sources
- NVIDIA BlogNVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut