From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

NVIDIA Blog News

Summary

NVIDIA explains how optimizing power management for AI factories using DSX Flex and partner platforms can boost efficiency, with Lambda validating a 24% throughput increase on a fixed power budget.

<div id="bsf_rt_marker"></div><p><span style="font-weight: 400;">On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked</span><span style="font-weight: 400;">, Silicon Valley Power</span><span style="font-weight: 400;"> sent a signal to an AI factory to adjust its power consumption.</span></p> <p><span style="font-weight: 400;">Varun Sivaram was watching on Zoom with about forty others — his team at </span><span style="font-weight: 400;">Emerald AI </span><span style="font-weight: 400;">in their San Francisco conference room, engineers at the data center and people from the utility itself. Nobody touched anything.</span></p> <p><span style="font-weight: 400;">Emerald AI’</span><span style="font-weight: 400;">s Conductor platform — a grid-orchestration platform from NVIDIA partner </span><span style="font-weight: 400;">Emerald AI,</span><span style="font-weight: 400;"> and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver  — receives signals about grid conditions and adjusts the data center’s flexible computing workloads. Work that can wait is slowed or rescheduled, while higher-priority services continue operating. </span></p> <p><span style="font-weight: 400;">The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly what </span><span style="font-weight: 400;">Silicon Valley Powe</span><span style="font-weight: 400;">r needed, </span></p> <p><span style="font-weight: 400;">When the reduction showed on screen, everyone cheered.</span></p> <p><span style="font-weight: 400;">“We were watching with bated breath,” Sivaram said. “It was our first time deploying across thousands of NVIDIA GPUs.” His head of product, Mansi Shah, was emotional. “This feels kind of like a SpaceX rocket launch,” she said.</span></p> <p><span style="font-weight: 400;">Silicon Valley Power h</span><span style="font-weight: 400;">as since sent more than 200 demand signals to that AI factory. It worked every single time. </span></p> <figure id="attachment_98197" aria-describedby="caption-attachment-98197" style="width: 1200px" class="wp-caption aligncenter"><img fetchpriority="high" decoding="async" class="size-large wp-image-98197" src="https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-1680x881.jpg" alt="" width="1200" height="629" srcset="https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-1680x881.jpg 1680w, https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-960x504.jpg 960w, https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-1280x671.jpg 1280w, https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-1536x806.jpg 1536w, https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-scaled.jpg 2048w, https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-630x330.jpg 630w" sizes="(max-width: 1200px) 100vw, 1200px" /><figcaption id="caption-attachment-98197" class="wp-caption-text">The Emerald AI team in San Francisco watches as Silicon Valley Power&#8217;s demand signal hits the factory floor — power dropping from four megawatts to three, automatically, while every high-priority job keeps running.</figcaption></figure> <p><span style="font-weight: 400;">This is grid flexibility in production. And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America’s AI factories need, without waiting a decade to build new transmission lines.</span></p> <p><span style="font-weight: 400;">At the</span><a target="_blank" href="https://www.ai-infra-summit.com/"> <span style="font-weight: 400;">AI Infra Summit</span></a><span style="font-weight: 400;"> on Tuesday, Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote. </span></p> <p><span style="font-weight: 400;">Results from cloud provider </span><span style="font-weight: 400;">Lambda’s </span><span style="font-weight: 400;">first validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently.</span></p> <p><span style="font-weight: 400;">“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services at </span><span style="font-weight: 400;">Lambda</span><span style="font-weight: 400;">. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint.” </span></p> <p><span style="font-weight: 400;">That August evening, when SVP called, Conductor executed against a predefined workload hierarchy: lowest-priority jobs yielded, high-priority inference kept running, and power fell from four megawatts to three. Automated. No operator required.</span><!-- DSX at a Glance — vest-pocket, float right --></p> <div style="float: right; width: 220px; margin: 0 0 24px 28px; clear: right;"> <aside style="padding: 18px 16px; background: #F5F5F5; border-right: 4px solid #76B900; font-family: 'NVIDIA Sans',Arial,sans-serif;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 14px; font-weight: bold; color: #000; margin: 0 0 14px 0; border-bottom: 1px solid #E0E0E0; padding-bottom: 10px;">DSX at a glance</div> <div style="margin-bottom: 12px;"> <div style="font-size: 12px; font-weight: bold; color: #000; margin-bottom: 4px; font-family: 'NVIDIA Sans',Arial,sans-serif;">NVIDIA DSX MaxLPS</div> <div style="font-size: 12px; line-height: 1.5; color: #313131; font-family: 'NVIDIA Sans',Arial,sans-serif;">A suite of technologies to optimize AI factory throughput per megawatt, including dynamic power allocation software that monitors GPU and rack-level consumption in real time, recovering stranded capacity to maximize token throughput within a fixed power budget.</div> </div> <div style="margin-bottom: 12px; border-top: 1px solid #E0E0E0; padding-top: 10px;"> <div style="font-size: 12px; font-weight: bold; color: #000; margin-bottom: 4px; font-family: 'NVIDIA Sans',Arial,sans-serif;">NVIDIA DSX Flex</div> <div style="font-size: 12px; line-height: 1.5; color: #313131; font-family: 'NVIDIA Sans',Arial,sans-serif;">Receives grid signals (load-shedding, demand-response, pricing events) and adapts AI workload priorities in response, protecting high-priority jobs while reducing overall power draw.</div> </div> <div style="margin-bottom: 12px; border-top: 1px solid #E0E0E0; padding-top: 10px;"> <div style="font-size: 12px; font-weight: bold; color: #000; margin-bottom: 4px; font-family: 'NVIDIA Sans',Arial,sans-serif;">NVIDIA DSX OS</div> <div style="font-size: 12px; line-height: 1.5; color: #313131; font-family: 'NVIDIA Sans',Arial,sans-serif;">Open-source, modular software for AI factory lifecycle management, runtime consistency, health automation and resiliency.</div> </div> <div style="margin-bottom: 12px; border-top: 1px solid #E0E0E0; padding-top: 10px;"> <div style="font-size: 12px; font-weight: bold; color: #000; margin-bottom: 4px; font-family: 'NVIDIA Sans',Arial,sans-serif;">NVIDIA DSX Sim</div> <div style="font-size: 12px; line-height: 1.5; color: #313131; font-family: 'NVIDIA Sans',Arial,sans-serif;">Simulation tools that let operators model and validate factory designs before physical deployment, identifying bottlenecks before capital is fixed.</div> </div> <div style="border-top: 1px solid #E0E0E0; padding-top: 10px;"> <div style="font-size: 12px; font-weight: bold; color: #000; margin-bottom: 4px; font-family: 'NVIDIA Sans',Arial,sans-serif;">NVIDIA DSX Reference Designs</div> <div style="font-size: 12px; line-height: 1.5; color: #313131; font-family: 'NVIDIA Sans',Arial,sans-serif;">Generation-specific, validated architectures spanning compute, networking, storage and facilities, co-designed with NVIDIA’s ecosystem partners.</div> </div> </aside> </div> <p><span style="font-weight: 400;">In the AI factory economy, power is the constraint. Work per gigawatt is the metric. Data center operators are meticulous about efficiency — every watt put to work is a watt delivering productive compute, and the industry has driven remarkable gains at every layer of the stack, from facility design to rack-level power conversion.</span></p> <p><span style="font-weight: 400;">DSX extends that discipline into the AI workload itself. Smarter rack provisioning puts power where workloads actually need it. Operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs running rather than waiting. The goal is the same one operators have always pursued: more work from the power you have.</span></p> <p><span style="font-weight: 400;">“A one-gigawatt factory will never become a two-gigawatt factory,” NVIDIA founder and CEO Jensen Huang has said.</span></p> <p><span style="font-weight: 400;">The answer engineers reach when systems hit physical limits is always the same: stop optimizing the parts and start designing the whole. </span></p> <p><span style="font-weight: 400;">Introduced at GTC Taipei in May, NVIDIA DSX is that answer for the AI factory — and the early deployments are already proving it out. </span></p> <p><span style="font-weight: 400;">The</span><a target="_blank" href="https://www.nvidia.com/en-us/data-center/products/dsx/"> <span style="font-weight: 400;">full platform</span></a><span style="font-weight: 400;"> spans networking, cooling, water efficiency and facility design; the sections below focus on some of the results so far in power management and grid participation.</span></p> <h2><span style="font-weight: 400;">More Compute, Same Budget: DSX MaxLPS</span></h2> <p><span style="font-weight: 400;">Lambda’s </span><span style="font-weight: 400;">results, released at the AI Infra Summit, are the first validation of DSX MaxLPS on</span><span style="font-weight: 400;"> NVIDIA HGX B20</span><span style="font-weight: 400;">0 GPU Servers. </span></p> <p><span style="font-weight: 400;">DSX MaxLPS monitors GPU and rack-level power consumption and reallocates headroom across nodes based on workload type, recovering capacity that static provisioning would leave stranded. Training and inference draw power differently; MaxLPS optimizes allocation in AI factories running both.</span></p> <p><span style="font-weight: 400;">Lambda,</span><span style="font-weight: 400;"> a GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers, ran the software on a five-rack, 19-node cluster. </span></p> <p><span style="font-weight: 400;">What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster-wide token throughput — from roughly 4 million tokens per second to 5 million. Performance per watt improved by 23%.</span></p> <p><span style="font-weight: 400;">Based on NVIDIA’s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments.</span></p> <h2><span style="font-weight: 400;">Automated Demand Response, Proven in Production</span></h2> <p><span style="font-weight: 400;">The Santa Clara story isn&#8217;t a DSX Flex installation — it&#8217;s something earlier and more important: proof that the concept works at commercial scale.</span>&lt;<br /> <span style="font-weight: 400;">NVIDIA&#8217;s Eos AI factory is running Emerald AI Conductor as a participant in </span><span style="font-weight: 400;">Silicon Valley Power</span><span style="font-weight: 400;">&#8216;s Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources. </span></p> <p><span style="font-weight: 400;">When </span><span style="font-weight: 400;">Silicon Valley Power</span><span style="font-weight: 400;"> sends a signal, Conductor responds in under a minute. The factory that&#8217;s willing to flex gets to run bigger.</span></p> <p><span style="font-weight: 400;">That&#8217;s the pattern DSX Flex is built to generalize — with Emerald AI Conductor integrating into DSX Flex as the platform matures. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia, facility: a 96-megawatt Vera Rubin AI factory at NVIDIA&#8217;s AI Factory Research Center, building on five prior demonstrations across two continents.</span></p> <h2><span style="font-weight: 400;">The Next Power Architecture Layer: 800V DC Power Architecture</span></h2> <p><span style="font-weight: 400;">The gains inside today&#8217;s AI factory are real and deployable now. The next layer is how power is delivered to denser accelerated computing racks.</span></p> <p>As AI factories scale, traditional lower-voltage power paths add conversion complexity and distribution constraints.</p> <p><span style="font-weight: 400;">NVIDIA’s 800 VDC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks.</span></p> <p><span style="font-weight: 400;">NVIDIA DSX is incorporating 800V DC into its reference designs. </span></p> <h2><span style="font-weight: 400;">The Whole Factory, Not the Parts</span></h2> <p><span style="font-weight: 400;">No single component can optimize an AI factory on its own. A faster GPU still waits on the network. Power can be stranded by bad provisioning. Cooling overhead still diverts electricity from GPUs; GB200 NVL72 racks running direct liquid cooling carry ~120 kW of heat that has to go somewhere before that power reaches compute.</span></p> <p><span style="font-weight: 400;">The only reliable path to more tokens per megawatt is to optimize the whole factory — DSX Sim before the first rack goes in, DSX OS and DSX Exchange once it&#8217;s running, DSX Reference Designs so builders start from a validated architecture rather than from scratch. (See sidebar for the full DSX suite at a glance.)</span></p> <h2><span style="font-weight: 400;">The Gigawatt Infrastructure Standard</span></h2> <p><span style="font-weight: 400;">It all comes down to one question: how much useful work does the factory produce per megawatt consumed? </span></p> <p><span style="font-weight: 400;">NVIDIA DSX gives infrastructure builders the reference designs, simulation tools, operational software, and power-management technology to compete on that metric, on current hardware and into the next generation.</span></p> <p><span style="font-weight: 400;">When the grid needed relief, the factory gave it without dropping a job, without asking for more power. </span></p> <p><span style="font-weight: 400;">With NVIDIA DSX, that&#8217;s the new baseline for what an AI factory is supposed to do.</span></p> <aside style="margin: 32px 0; padding: 24px 28px; background: #F5F5F5; border-left: 4px solid #76B900; font-family: 'NVIDIA Sans',Arial,sans-serif;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 18px; line-height: 1.3; color: #000000; font-weight: bold; margin: 0 0 18px 0;">Numbers at a glance</div> <div style="display: table; width: 100%; table-layout: fixed; border-bottom: 1px solid #E0E0E0; padding-bottom: 8px; margin-bottom: 4px;"> <div style="display: table-cell; width: 28%; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; font-weight: bold; color: #757575; text-transform: uppercase; letter-spacing: 0.06em; vertical-align: bottom;">Metric</div> <div style="display: table-cell; width: 48%; padding-left: 16px; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; font-weight: bold; color: #757575; text-transform: uppercase; letter-spacing: 0.06em; vertical-align: bottom;">Context</div> <div style="display: table-cell; width: 24%; padding-left: 12px; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; font-weight: bold; color: #757575; text-transform: uppercase; letter-spacing: 0.06em; vertical-align: bottom;">Source</div> </div> <div style="display: table; width: 100%; table-layout: fixed; padding: 12px 0; border-bottom: 1px solid #E0E0E0;"> <div style="display: table-cell; width: 28%; vertical-align: top;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 22px; line-height: 1; color: #76b900; font-weight: bold;">+24%</div> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; line-height: 1.4; color: #4b4b4b; font-weight: 500; margin-top: 4px;">cluster token throughput</div> </div> <div style="display: table-cell; width: 48%; padding-left: 16px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 13px; line-height: 1.5; color: #313131;">19 nodes at 85% power vs. 16 nodes at full power, same facility budget</div> <div style="display: table-cell; width: 24%; padding-left: 12px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 12px; line-height: 1.45; color: #757575;">Lambda, HGX B200, DSX MaxLPS</div> </div> <div style="display: table; width: 100%; table-layout: fixed; padding: 12px 0; border-bottom: 1px solid #E0E0E0;"> <div style="display: table-cell; width: 28%; vertical-align: top;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 22px; line-height: 1; color: #76b900; font-weight: bold;">+23%</div> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; line-height: 1.4; color: #4b4b4b; font-weight: 500; margin-top: 4px;">performance per watt</div> </div> <div style="display: table-cell; width: 48%; padding-left: 16px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 13px; line-height: 1.5; color: #313131;">19-node cluster at 85% power policy vs. 16-node full-power baseline</div> <div style="display: table-cell; width: 24%; padding-left: 12px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 12px; line-height: 1.45; color: #757575;">Lambda, HGX B200, DSX MaxLPS</div> </div> <div style="display: table; width: 100%; table-layout: fixed; padding: 12px 0; border-bottom: 1px solid #E0E0E0;"> <div style="display: table-cell; width: 28%; vertical-align: top;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 22px; line-height: 1; color: #76b900; font-weight: bold;">40%</div> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; line-height: 1.4; color: #4b4b4b; font-weight: 500; margin-top: 4px;">power demand reduction in under a minute</div> </div> <div style="display: table-cell; width: 48%; padding-left: 16px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 13px; line-height: 1.5; color: #313131;">SVP automated response, Flexible Load Interconnect Program</div> <div style="display: table-cell; width: 24%; padding-left: 12px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 12px; line-height: 1.45; color: #757575;">Emerald AI / Future DSX Flex</div> </div> <div style="display: table; width: 100%; table-layout: fixed; padding: 12px 0; border-bottom: 1px solid #E0E0E0;"> <div style="display: table-cell; width: 28%; vertical-align: top;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 22px; line-height: 1; color: #76b900; font-weight: bold;">Up to +40%</div> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; line-height: 1.4; color: #4b4b4b; font-weight: 500; margin-top: 4px;">more GPU capacity</div> </div> <div style="display: table-cell; width: 48%; padding-left: 16px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 13px; line-height: 1.5; color: #313131;">Vera Rubin NVL72, MaxLPS combined with data center power planning, same power budget</div> <div style="display: table-cell; width: 24%; padding-left: 12px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 12px; line-height: 1.45; color: #757575;">DSX MaxLPS</div> </div> <div style="display: table; width: 100%; table-layout: fixed; padding: 12px 0;"> <div style="display: table-cell; width: 28%; vertical-align: top;"> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 22px; line-height: 1; color: #76b900; font-weight: bold;">3–5%</div> <div style="font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 11px; line-height: 1.4; color: #4b4b4b; font-weight: 500; margin-top: 4px;">end-to-end efficiency gain (projected)</div> </div> <div style="display: table-cell; width: 48%; padding-left: 16px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 13px; line-height: 1.5; color: #313131;">800V DC vs. 54V distribution; available with Vera Rubin NVL72 2027</div> <div style="display: table-cell; width: 24%; padding-left: 12px; vertical-align: top; font-family: 'NVIDIA Sans',Arial,sans-serif; font-size: 12px; line-height: 1.45; color: #757575;">NVIDIA DSX 800V DC architecture</div> </div> </aside>
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:31 PM

# From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production Source: [https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/](https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/) On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Powersent a signal to an AI factory to adjust its power consumption\. Varun Sivaram was watching on Zoom with about forty others — his team atEmerald AIin their San Francisco conference room, engineers at the data center and people from the utility itself\. Nobody touched anything\. Emerald AI’s Conductor platform — a grid\-orchestration platform from NVIDIA partnerEmerald AI,and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver — receives signals about grid conditions and adjusts the data center’s flexible computing workloads\. Work that can wait is slowed or rescheduled, while higher\-priority services continue operating\. The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly whatSilicon Valley Power needed, When the reduction showed on screen, everyone cheered\. “We were watching with bated breath,” Sivaram said\. “It was our first time deploying across thousands of NVIDIA GPUs\.” His head of product, Mansi Shah, was emotional\. “This feels kind of like a SpaceX rocket launch,” she said\. Silicon Valley Power has since sent more than 200 demand signals to that AI factory\. It worked every single time\. ![](https://blogs.nvidia.com/wp-content/uploads/3026/09/Nvidia-Image-2-1680x881.jpg)The Emerald AI team in San Francisco watches as Silicon Valley Power’s demand signal hits the factory floor — power dropping from four megawatts to three, automatically, while every high\-priority job keeps running\.This is grid flexibility in production\. And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America’s AI factories need, without waiting a decade to build new transmission lines\. At the[AI Infra Summit](https://www.ai-infra-summit.com/)on Tuesday, Ian Buck, NVIDIA’s vice president of hyperscale and high\-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote\. Results from cloud providerLambda’sfirst validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently\. “With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services atLambda\. “NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real\-world usage, with significantly more compute density in the same footprint\.” That August evening, when SVP called, Conductor executed against a predefined workload hierarchy: lowest\-priority jobs yielded, high\-priority inference kept running, and power fell from four megawatts to three\. Automated\. No operator required\. In the AI factory economy, power is the constraint\. Work per gigawatt is the metric\. Data center operators are meticulous about efficiency — every watt put to work is a watt delivering productive compute, and the industry has driven remarkable gains at every layer of the stack, from facility design to rack\-level power conversion\. DSX extends that discipline into the AI workload itself\. Smarter rack provisioning puts power where workloads actually need it\. Operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs running rather than waiting\. The goal is the same one operators have always pursued: more work from the power you have\. “A one\-gigawatt factory will never become a two\-gigawatt factory,” NVIDIA founder and CEO Jensen Huang has said\. The answer engineers reach when systems hit physical limits is always the same: stop optimizing the parts and start designing the whole\. Introduced at GTC Taipei in May, NVIDIA DSX is that answer for the AI factory — and the early deployments are already proving it out\. The[full platform](https://www.nvidia.com/en-us/data-center/products/dsx/)spans networking, cooling, water efficiency and facility design; the sections below focus on some of the results so far in power management and grid participation\. ## More Compute, Same Budget: DSX MaxLPS Lambda’sresults, released at the AI Infra Summit, are the first validation of DSX MaxLPS onNVIDIA HGX B200 GPU Servers\. DSX MaxLPS monitors GPU and rack\-level power consumption and reallocates headroom across nodes based on workload type, recovering capacity that static provisioning would leave stranded\. Training and inference draw power differently; MaxLPS optimizes allocation in AI factories running both\. Lambda,a GPU cloud provider serving more than 10,000 customers from AI\-native startups to hyperscalers, ran the software on a five\-rack, 19\-node cluster\. What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster\-wide token throughput — from roughly 4 million tokens per second to 5 million\. Performance per watt improved by 23%\. Based on NVIDIA’s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next\-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments\. ## Automated Demand Response, Proven in Production The Santa Clara story isn’t a DSX Flex installation — it’s something earlier and more important: proof that the concept works at commercial scale\.< NVIDIA’s Eos AI factory is running Emerald AI Conductor as a participant inSilicon Valley Power‘s Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources\. WhenSilicon Valley Powersends a signal, Conductor responds in under a minute\. The factory that’s willing to flex gets to run bigger\. That’s the pattern DSX Flex is built to generalize — with Emerald AI Conductor integrating into DSX Flex as the platform matures\. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia, facility: a 96\-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center, building on five prior demonstrations across two continents\. ## The Next Power Architecture Layer: 800V DC Power Architecture The gains inside today’s AI factory are real and deployable now\. The next layer is how power is delivered to denser accelerated computing racks\. As AI factories scale, traditional lower\-voltage power paths add conversion complexity and distribution constraints\. NVIDIA’s 800 VDC architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks\. NVIDIA DSX is incorporating 800V DC into its reference designs\. ## The Whole Factory, Not the Parts No single component can optimize an AI factory on its own\. A faster GPU still waits on the network\. Power can be stranded by bad provisioning\. Cooling overhead still diverts electricity from GPUs; GB200 NVL72 racks running direct liquid cooling carry ~120 kW of heat that has to go somewhere before that power reaches compute\. The only reliable path to more tokens per megawatt is to optimize the whole factory — DSX Sim before the first rack goes in, DSX OS and DSX Exchange once it’s running, DSX Reference Designs so builders start from a validated architecture rather than from scratch\. \(See sidebar for the full DSX suite at a glance\.\) ## The Gigawatt Infrastructure Standard It all comes down to one question: how much useful work does the factory produce per megawatt consumed? NVIDIA DSX gives infrastructure builders the reference designs, simulation tools, operational software, and power\-management technology to compete on that metric, on current hardware and into the next generation\. When the grid needed relief, the factory gave it without dropping a job, without asking for more power\. With NVIDIA DSX, that’s the new baseline for what an AI factory is supposed to do\.

Similar Articles