{"id":4687,"date":"2026-09-16T18:40:18","date_gmt":"2026-09-16T18:40:18","guid":{"rendered":"https:\/\/salarydistribution.com\/machine-learning\/2026\/09\/16\/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut\/"},"modified":"2026-09-16T18:40:18","modified_gmt":"2026-09-16T18:40:18","slug":"nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut","status":"publish","type":"post","link":"https:\/\/salarydistribution.com\/machine-learning\/2026\/09\/16\/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut\/","title":{"rendered":"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut"},"content":{"rendered":"<div>\n<p><span>System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments.\u00a0<\/span><\/p>\n<p><span>Underlying all three is platform fungibility: the same infrastructure runs any model, any workload, from training to inference, recommender to reasoning, language to video, keeping utilization high.<\/span><\/p>\n<p><span>The NVIDIA platform is purpose-built to optimize across all these, as highlighted by MLPerf Inference v6.1 results released today:<\/span><\/p>\n<ul>\n<li><b>NVIDIA Vera Rubin NVL72 system debuts with leading performance<\/b><span>: In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to <\/span><span>3.7x<\/span><span> better throughput than GB300 NVL72.<\/span><\/li>\n<li><b>NVIDIA GB300 NVL72 scales<\/b> <b>with leading efficiency<\/b><span>: A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.<\/span><\/li>\n<li><b>Continuous software optimizations drive performance gains<\/b><span>: Software optimizations in NVIDIA\u2019s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0. Optimizations continued post-v6.1 submission, delivering further performance gains.<\/span><\/li>\n<\/ul>\n<p><span>For organizations making AI infrastructure decisions, performance, scaling efficiency and software velocity are important considerations that determine long-term inference economics.\u00a0<\/span><\/p>\n<h2><b>Vera Rubin NVL72 Makes MLPerf Inference Debut With Leading Performance<\/b><\/h2>\n<p><span>NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.\u00a0<\/span><\/p>\n<p><span>Vera Rubin NVL72 delivers up to <\/span><span>3.7x<\/span><span> higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72. These early results showcase NVIDIA\u2019s accelerated pace of innovation and how performance will improve with continuous software optimizations.\u00a0<\/span><\/p>\n<figure id=\"attachment_98218\" aria-describedby=\"caption-attachment-98218\" class=\"wp-caption alignnone\"><img decoding=\"async\" loading=\"lazy\" class=\"wp-image-98218 size-full\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/nvidia-vera-rubin-delivers-up-to-3-7x-better-performance.jpeg\" alt=\"\" width=\"1280\" height=\"720\"><figcaption id=\"caption-attachment-98218\" class=\"wp-caption-text\">MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0106 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.<\/figcaption><\/figure>\n<p><span>This performance means each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack, while lowering cost per token.<\/span><\/p>\n<p><span>The results reflect full-stack codesign across hardware and software. Vera Rubin\u2019s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces memory footprint across model weights, attention and KV cache \u2014 increasing throughput with minimal loss of output quality.\u00a0<\/span><\/p>\n<p><span>Vera Rubin submissions heavily used disaggregated serving, separating prefill and decode along with large-scale expert parallelism for maximum efficiency across the <\/span><a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-us\/glossary\/mixture-of-experts\/\" rel=\"noopener\"><span>mixture-of-experts<\/span><\/a><span> layers that power models like DeepSeek-R1 and Qwen3-VL.\u00a0<\/span><\/p>\n<p><span>The NVL72 scale-up domain \u2014 powered by sixth-generation NVIDIA NVLink and NVLink Switch to deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet \u2014 provides the interconnect foundation that makes these techniques effective at rack scale.\u00a0<\/span><\/p>\n<p><span>This codesign extends to NVIDIA\u2019s partner ecosystem: <\/span><a target=\"_blank\" href=\"https:\/\/nebius.com\/blog\/posts\/mlperf-inference-v6-1-results\" rel=\"noopener\"><span>Nebius <\/span><\/a><span>also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.<\/span><\/p>\n<p><span>AI agents, which reason, plan and act across multiple steps, are reshaping how inference performance is measured. In benchmarks designed to capture this shift, such as SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30<\/span><span>x <\/span><span>better performance than GB300 NVL72 in preview testing. In addition, the upcoming MLPerf Endpoints benchmark will bring standardized measurement to agentic inference workloads, beyond what traditional throughput benchmarks capture.<\/span><\/p>\n<h2><b>NVIDIA GB300 NVL72 Scales With Leading Efficiency<\/b><\/h2>\n<p><span>Scaling efficiency \u2014 how effectively additional GPUs translate to throughput gains \u2014 is a key measure of AI infrastructure productivity. NVIDIA delivers this with high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks and efficient request orchestration across nodes.<\/span><\/p>\n<p><span>NVIDIA\u2019s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. Throughput grew nearly in proportion to the hardware added.<\/span><\/p>\n<figure id=\"attachment_98217\" aria-describedby=\"caption-attachment-98217\" class=\"wp-caption alignnone\"><img decoding=\"async\" loading=\"lazy\" class=\"size-full wp-image-98217\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/nvidia-gb300-nvl72-scaling-efficiency.jpeg\" alt=\"\" width=\"1280\" height=\"720\"><figcaption id=\"caption-attachment-98217\" class=\"wp-caption-text\">MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.<\/figcaption><\/figure>\n<p><span>Scaling efficiency is key because more GPUs don\u2019t automatically mean proportionally more throughput. If adding nearly double the GPU count delivered only a single-digit percentage improvement in throughput, the infrastructure cost would far outpace the performance return. The architecture, interconnect and software must all scale together.<\/span><\/p>\n<p><span>GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video \u2014 9x higher throughput and 7.5x lower latency than a single node.<\/span><\/p>\n<h2><b>Software Optimizations Drive Continuous Gains<\/b><\/h2>\n<p><span>NVIDIA platform undergoes continuous software development, delivering performance and feature improvements.\u00a0<\/span><\/p>\n<p><span>In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results. The gains came through lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and NVIDIA Dynamo.\u00a0<\/span><\/p>\n<p><span>Software optimization continued past the v6.1 submission deadline as well. Post-submission results, not yet verified by MLCommons, on GPT-OSS-120B and DLRMv3 show further performance gains.\u00a0<\/span><\/p>\n<p><img decoding=\"async\" loading=\"lazy\" class=\"alignnone size-full wp-image-98216\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/nvidia-delivers-continuous-software-optimizations.jpeg\" alt=\"\" width=\"1280\" height=\"720\"><\/p>\n<h2><b>AI Inference at Every Scale<\/b><\/h2>\n<p><span>Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B.<\/span><\/p>\n<p><span>The NVIDIA partner ecosystem participated broadly, with 19 partners \u2014 eight of them on multi-node Blackwell NVL72 systems \u2014 demonstrating excellent performance. This includes <\/span><span>ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro and Wiwynn<\/span><span>.<\/span><\/p>\n<p><span>From compact edge devices to the largest AI factories, NVIDIA continues to advance performance across the full technology stack with an annual cadence of platform architectures, continuously improving software and an ecosystem built to deliver it at scale.<\/span><\/p>\n<p><i><span>Learn more about the <\/span><\/i><a href=\"https:\/\/blogs.nvidia.com\/blog\/vera-rubin\/\"><i><span>NVIDIA Vera Rubin platform<\/span><\/i><\/a><i><span>.<\/span><\/i><\/p>\n<\/p><\/div>\n","protected":false},"excerpt":{"rendered":"<p>https:\/\/blogs.nvidia.com\/blog\/vera-rubin-nvl72-mlperf-inference\/<\/p>\n","protected":false},"author":0,"featured_media":4688,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[3],"tags":[],"_links":{"self":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts\/4687"}],"collection":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/comments?post=4687"}],"version-history":[{"count":0,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts\/4687\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/media\/4688"}],"wp:attachment":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/media?parent=4687"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/categories?post=4687"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/tags?post=4687"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}