THE CRUNCH

NVIDIA has unveiled its Vera Rubin NVL72 system, which debuted in the MLPerf Inference v6.1 benchmarks with leading performance. The system delivers up to 3.7x higher throughput than the previous GB300 NVL72 on the Qwen3-VL benchmark and up to 2.5x higher throughput on DeepSeek-R1. These gains are attributed to a full-stack codesign that includes enhanced Tensor Cores, Transformer Engine acceleration, and the use of

NVFP4 precision to reduce memory footprint. The system also uses disaggregated serving and large-scale expert parallelism to maximise efficiency across mixture-of-experts layers. In addition to raw throughput, the GB300 NVL72 achieved 99% scaling efficiency when scaled from a single rack to four racks, indicating that throughput grows nearly linearly with hardware additions. NVIDIA also highlighted that continuous software optimisations delivered up to 1.6x higher performance over the previous v6.0 benchmark version.

WHAT HAPPENS NEXT

NVIDIA plans to continue software optimisations post-v6.1 submission, which will likely deliver further performance gains. The company also noted that the upcoming MLPerf Endpoints benchmark will bring standardised measurement to agentic inference workloads.