{"id":4681,"date":"2026-09-15T20:40:04","date_gmt":"2026-09-15T20:40:04","guid":{"rendered":"https:\/\/salarydistribution.com\/machine-learning\/2026\/09\/15\/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production\/"},"modified":"2026-09-15T20:40:04","modified_gmt":"2026-09-15T20:40:04","slug":"from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production","status":"publish","type":"post","link":"https:\/\/salarydistribution.com\/machine-learning\/2026\/09\/15\/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production\/","title":{"rendered":"From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production"},"content":{"rendered":"<div>\n<p><span>On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked<\/span><span>, Silicon Valley Power<\/span><span> sent a signal to an AI factory to adjust its power consumption.<\/span><\/p>\n<p><span>Varun Sivaram was watching on Zoom with about forty others \u2014 his team at <\/span><span>Emerald AI <\/span><span>in their San Francisco conference room, engineers at the data center and people from the utility itself. Nobody touched anything.<\/span><\/p>\n<p><span>Emerald AI\u2019<\/span><span>s Conductor platform \u2014 a grid-orchestration platform from NVIDIA partner <\/span><span>Emerald AI,<\/span><span> and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver\u00a0 \u2014 receives signals about grid conditions and adjusts the data center\u2019s flexible computing workloads. Work that can wait is slowed or rescheduled, while higher-priority services continue operating.\u00a0<\/span><\/p>\n<p><span>The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads\u2014 exactly what <\/span><span>Silicon Valley Powe<\/span><span>r needed, <\/span><\/p>\n<p><span>When the reduction showed on screen, everyone cheered.<\/span><\/p>\n<p><span>\u201cWe were watching with bated breath,\u201d Sivaram said. \u201cIt was our first time deploying across thousands of NVIDIA GPUs.\u201d His head of product, Mansi Shah, was emotional. \u201cThis feels kind of like a SpaceX rocket launch,\u201d she said.<\/span><\/p>\n<p><span>Silicon Valley Power h<\/span><span>as since sent more than 200 demand signals to that AI factory. It worked every single time.\u00a0<\/span><\/p>\n<figure id=\"attachment_98197\" aria-describedby=\"caption-attachment-98197\" class=\"wp-caption aligncenter\"><img decoding=\"async\" loading=\"lazy\" class=\"size-large wp-image-98197\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/3026\/09\/Nvidia-Image-2-1680x881.jpg\" alt=\"\" width=\"1200\" height=\"629\"><figcaption id=\"caption-attachment-98197\" class=\"wp-caption-text\">The Emerald AI team in San Francisco watches as Silicon Valley Power\u2019s demand signal hits the factory floor \u2014 power dropping from four megawatts to three, automatically, while every high-priority job keeps running.<\/figcaption><\/figure>\n<p><span>This is grid flexibility in production. And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America\u2019s AI factories need, without waiting a decade to build new transmission lines.<\/span><\/p>\n<p><span>At the<\/span><a target=\"_blank\" href=\"https:\/\/www.ai-infra-summit.com\/\" rel=\"noopener\"> <span>AI Infra Summit<\/span><\/a><span> on Tuesday, Ian Buck, NVIDIA\u2019s vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote.\u00a0<\/span><\/p>\n<p><span>Results from cloud provider <\/span><span>Lambda\u2019s <\/span><span>first validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently.<\/span><\/p>\n<p><span>\u201cWith our proof of concept, we believe we\u2019ve moved beyond the limitation of fixed power budgets,\u201d said Dave Ward, president of cloud services at <\/span><span>Lambda<\/span><span>. \u201cNVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.\u201d\u00a0<\/span><\/p>\n<p><span>That August evening, when SVP called, Conductor executed against a predefined workload hierarchy: lowest-priority jobs yielded, high-priority inference kept running, and power fell from four megawatts to three. Automated. No operator required.<\/span><!-- DSX at a Glance &mdash; vest-pocket, float right --><\/p>\n<div>\n<aside>\n<p>DSX at a glance<\/p>\n<div>\n<p>NVIDIA DSX MaxLPS<\/p>\n<p>A suite of technologies to optimize AI factory throughput per megawatt, including dynamic power allocation software that monitors GPU and rack-level consumption in real time, recovering stranded capacity to maximize token throughput within a fixed power budget.<\/p>\n<\/div>\n<div>\n<p>NVIDIA DSX Flex<\/p>\n<p>Receives grid signals (load-shedding, demand-response, pricing events) and adapts AI workload priorities in response, protecting high-priority jobs while reducing overall power draw.<\/p>\n<\/div>\n<div>\n<p>NVIDIA DSX OS<\/p>\n<p>Open-source, modular software for AI factory lifecycle management, runtime consistency, health automation and resiliency.<\/p>\n<\/div>\n<div>\n<p>NVIDIA DSX Sim<\/p>\n<p>Simulation tools that let operators model and validate factory designs before physical deployment, identifying bottlenecks before capital is fixed.<\/p>\n<\/div>\n<div>\n<p>NVIDIA DSX Reference Designs<\/p>\n<p>Generation-specific, validated architectures spanning compute, networking, storage and facilities, co-designed with NVIDIA\u2019s ecosystem partners.<\/p>\n<\/div>\n<\/aside>\n<\/div>\n<p><span>In the AI factory economy, power is the constraint. Work per gigawatt is the metric. Data center operators are meticulous about efficiency \u2014 every watt put to work is a watt delivering productive compute, and the industry has driven remarkable gains at every layer of the stack, from facility design to rack-level power conversion.<\/span><\/p>\n<p><span>DSX extends that discipline into the AI workload itself. Smarter rack provisioning puts power where workloads actually need it. Operational intelligence \u2014 tighter scheduling, faster restarts, leaner checkpointing \u2014 keeps GPUs running rather than waiting. The goal is the same one operators have always pursued: more work from the power you have.<\/span><\/p>\n<p><span>\u201cA one-gigawatt factory will never become a two-gigawatt factory,\u201d NVIDIA founder and CEO Jensen Huang has said.<\/span><\/p>\n<p><span>The answer engineers reach when systems hit physical limits is always the same: stop optimizing the parts and start designing the whole.\u00a0<\/span><\/p>\n<p><span>Introduced at GTC Taipei in May, NVIDIA DSX is that answer for the AI factory \u2014 and the early deployments are already proving it out.\u00a0<\/span><\/p>\n<p><span>The<\/span><a target=\"_blank\" href=\"https:\/\/www.nvidia.com\/en-us\/data-center\/products\/dsx\/\" rel=\"noopener\"> <span>full platform<\/span><\/a><span> spans networking, cooling, water efficiency and facility design; the sections below focus on some of the results so far in power management and grid participation.<\/span><\/p>\n<h2><span>More Compute, Same Budget: DSX MaxLPS<\/span><\/h2>\n<p><span>Lambda\u2019s <\/span><span>results, released at the AI Infra Summit, are the first validation of DSX MaxLPS on<\/span><span> NVIDIA HGX B20<\/span><span>0 GPU Servers.\u00a0<\/span><\/p>\n<p><span>DSX MaxLPS monitors GPU and rack-level power consumption and reallocates headroom across nodes based on workload type, recovering capacity that static provisioning would leave stranded. Training and inference draw power differently; MaxLPS optimizes allocation in AI factories running both.<\/span><\/p>\n<p><span>Lambda,<\/span><span> a GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers, ran the software on a five-rack, 19-node cluster.\u00a0<\/span><\/p>\n<p><span>What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster-wide token throughput \u2014 from roughly 4 million tokens per second to 5 million. Performance per watt improved by 23%.<\/span><\/p>\n<p><span>Based on NVIDIA\u2019s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments.<\/span><\/p>\n<h2><span>Automated Demand Response, Proven in Production<\/span><\/h2>\n<p><span>The Santa Clara story isn\u2019t a DSX Flex installation \u2014 it\u2019s something earlier and more important: proof that the concept works at commercial scale.<\/span>&lt;<br \/><span>NVIDIA\u2019s Eos AI factory is running Emerald AI Conductor as a participant in <\/span><span>Silicon Valley Power<\/span><span>\u2018s Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources.\u00a0<\/span><\/p>\n<p><span>When <\/span><span>Silicon Valley Power<\/span><span> sends a signal, Conductor responds in under a minute. The factory that\u2019s willing to flex gets to run bigger.<\/span><\/p>\n<p><span>That\u2019s the pattern DSX Flex is built to generalize \u2014 with Emerald AI Conductor integrating into DSX Flex as the platform matures. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia, facility: a 96-megawatt Vera Rubin AI factory at NVIDIA\u2019s AI Factory Research Center, building on five prior demonstrations across two continents.<\/span><\/p>\n<h2><span>The Next Power Architecture Layer: 800V DC Power Architecture<\/span><\/h2>\n<p><span>The gains inside today\u2019s AI factory are real and deployable now. The next layer is how power is delivered to denser accelerated computing racks.<\/span><\/p>\n<p>As AI factories scale, traditional lower-voltage power paths add conversion complexity and distribution constraints.<\/p>\n<p><span>NVIDIA\u2019s 800 VDC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks.<\/span><\/p>\n<p><span>NVIDIA DSX is incorporating 800V DC into its reference designs.\u00a0<\/span><\/p>\n<h2><span>The Whole Factory, Not the Parts<\/span><\/h2>\n<p><span>No single component can optimize an AI factory on its own. A faster GPU still waits on the network. Power can be stranded by bad provisioning. Cooling overhead still diverts electricity from GPUs; GB200 NVL72 racks running direct liquid cooling carry ~120 kW of heat that has to go somewhere before that power reaches compute.<\/span><\/p>\n<p><span>The only reliable path to more tokens per megawatt is to optimize the whole factory \u2014 DSX Sim before the first rack goes in, DSX OS and DSX Exchange once it\u2019s running, DSX Reference Designs so builders start from a validated architecture rather than from scratch. (See sidebar for the full DSX suite at a glance.)<\/span><\/p>\n<h2><span>The Gigawatt Infrastructure Standard<\/span><\/h2>\n<p><span>It all comes down to one question: how much useful work does the factory produce per megawatt consumed?\u00a0<\/span><\/p>\n<p><span>NVIDIA DSX gives infrastructure builders the reference designs, simulation tools, operational software, and power-management technology to compete on that metric, on current hardware and into the next generation.<\/span><\/p>\n<p><span>When the grid needed relief, the factory gave it without dropping a job, without asking for more power.\u00a0<\/span><\/p>\n<p><span>With NVIDIA DSX, that\u2019s the new baseline for what an AI factory is supposed to do.<\/span><\/p>\n<aside>\n<p>Numbers at a glance<\/p>\n<div>\n<div>\n<p>+24%<\/p>\n<p>cluster token throughput<\/p>\n<\/div>\n<p>19 nodes at 85% power vs. 16 nodes at full power, same facility budget<\/p>\n<p>Lambda, HGX B200, DSX MaxLPS<\/p>\n<\/div>\n<div>\n<div>\n<p>+23%<\/p>\n<p>performance per watt<\/p>\n<\/div>\n<p>19-node cluster at 85% power policy vs. 16-node full-power baseline<\/p>\n<p>Lambda, HGX B200, DSX MaxLPS<\/p>\n<\/div>\n<div>\n<div>\n<p>40%<\/p>\n<p>power demand reduction in under a minute<\/p>\n<\/div>\n<p>SVP automated response, Flexible Load Interconnect Program<\/p>\n<p>Emerald AI \/ Future DSX Flex<\/p>\n<\/div>\n<div>\n<div>\n<p>Up to +40%<\/p>\n<p>more GPU capacity<\/p>\n<\/div>\n<p>Vera Rubin NVL72, MaxLPS combined with data center power planning, same power budget<\/p>\n<p>DSX MaxLPS<\/p>\n<\/div>\n<div>\n<div>\n<p>3\u20135%<\/p>\n<p>end-to-end efficiency gain (projected)<\/p>\n<\/div>\n<p>800V DC vs. 54V distribution; available with Vera Rubin NVL72 2027<\/p>\n<p>NVIDIA DSX 800V DC architecture<\/p>\n<\/div>\n<\/aside><\/div>\n","protected":false},"excerpt":{"rendered":"<p>https:\/\/blogs.nvidia.com\/blog\/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production\/<\/p>\n","protected":false},"author":0,"featured_media":4682,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[3],"tags":[],"_links":{"self":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts\/4681"}],"collection":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/comments?post=4681"}],"version-history":[{"count":0,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/posts\/4681\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/media\/4682"}],"wp:attachment":[{"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/media?parent=4681"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/categories?post=4681"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/salarydistribution.com\/machine-learning\/wp-json\/wp\/v2\/tags?post=4681"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}