NVIDIA announces that Groq 3 LPX is now in full production, delivering ultrafast token generation for agentic AI systems with the Vera Rubin platform, showing 4x faster performance in benchmarks.
<div id="bsf_rt_marker"></div><p><span style="font-weight: 400;">The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.</span></p>
<p><a target="_blank" href="https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai"><span style="font-weight: 400;">Announced today</span></a><span style="font-weight: 400;">, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform. </span></p>
<p><span style="font-weight: 400;">Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.</span></p>
<p><span style="font-weight: 400;">As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems. </span></p>
<p><span style="font-weight: 400;">These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale.</span><a target="_blank" href="https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/Rubin%20CPX%20Product%20Aug%202026.pdf"><span style="font-weight: 400;"> </span></a></p>
<p><span style="font-weight: 400;">At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.</span></p>
<h3><b>Extreme Codesign Optimizes for Performance</b></h3>
<p><span style="font-weight: 400;">Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences. </span></p>
<p><a target="_blank" href="https://www.nvidia.com/en-us/networking/spectrumx/"><span style="font-weight: 400;">NVIDIA Spectrum-X Ethernet</span></a><span style="font-weight: 400;"> moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.</span></p>
<p><span style="font-weight: 400;">NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.</span></p>
<p><span style="font-weight: 400;">Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue. </span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#nvidia-partners"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="nvidia-partners" class="wp-block-heading" style="scroll-margin-top: 100px;">NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs</h2>
<p><span style="font-weight: 400;">Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications. </span></p>
<p><span style="font-weight: 400;">Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.</span></p>
<p><span style="font-weight: 400;">Connecting NVIDIA Vera Rubin racks, </span><span style="font-weight: 400;">CoreWeave</span><span style="font-weight: 400;"> is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure. </span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#spacexai-vera-cpus"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="spacexai-vera-cpus" class="wp-block-heading" style="scroll-margin-top: 100px;"><b>SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI</b></h2>
<p><span style="font-weight: 400;">SpaceXAI </span><span style="font-weight: 400;">plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation. </span></p>
<p><img decoding="async" class="alignnone size-medium wp-image-97839" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-960x540.jpg" alt="" width="960" height="540" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-960x540.jpg 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-1680x945.jpg 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-1280x720.jpg 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-1536x864.jpg 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-1290x725.jpg 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-630x354.jpg 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-300x169.jpg 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750-400x225.jpg 400w, https://blogs.nvidia.com/wp-content/uploads/2026/08/cpu-press-hot-chips-vera-cpu-1920x1080-5562750.jpg 1920w" sizes="(max-width: 960px) 100vw, 960px" /></p>
<p><span style="font-weight: 400;">The SpaceXAI </span><span style="font-weight: 400;"><a target="_blank" href="https://nvidianews.nvidia.com/news/spacexai-adopts-nvidia-vera-cpu-to-accelerate-agentic-ai-at-massive-scale">partnership</a> extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.</span></p>
<p><span style="font-weight: 400;">Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized. </span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#groq-3-lpx"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="groq-3-lpx" class="wp-block-heading" style="scroll-margin-top: 100px;"><b>NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator</b></h2>
<p><span style="font-weight: 400;">Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.</span></p>
<p><span style="font-weight: 400;">Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation. </span></p>
<p><span style="font-weight: 400;">NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency. </span></p>
<p><img decoding="async" class="alignnone size-medium wp-image-97900" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-960x540.png" alt="" width="960" height="540" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-1680x944.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-scaled.png 2048w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/LPX-Chart-400x225.png 400w" sizes="(max-width: 960px) 100vw, 960px" /></p>
<p><span style="font-weight: 400;">Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.</span></p>
<h3><b>Building the Token Factory</b></h3>
<p><span style="font-weight: 400;">As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.</span></p>
<p><span style="font-weight: 400;">NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization.</span><a target="_blank" href="https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/LPXDecks.pdf"><span style="font-weight: 400;"> </span></a></p>
<h3><b>Extreme Codesign for Inference</b></h3>
<p><span style="font-weight: 400;">Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads. </span></p>
<p><span style="font-weight: 400;">At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.</span></p>
<h3><b>Designed for the Agentic AI Era</b></h3>
<p><span style="font-weight: 400;">As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.</span></p>
<p><span style="font-weight: 400;">And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned. </span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#spectrum-x-multiplane"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="spectrum-x-multiplane" class="wp-block-heading" style="scroll-margin-top: 100px;">NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network</h2>
<p><span style="font-weight: 400;">As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting </span><a target="_blank" href="https://www.nvidia.com/en-us/networking/spectrumx/"><span style="font-weight: 400;">Spectrum-X Multiplane</span></a><span style="font-weight: 400;"> — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.</span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-medium wp-image-97841" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-960x540.png" alt="" width="960" height="540" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-960x540.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-1680x945.png 1680w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-1280x720.png 1280w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-1536x864.png 1536w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-scaled.png 2048w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-1290x725.png 1290w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-630x354.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-300x169.png 300w, https://blogs.nvidia.com/wp-content/uploads/2026/08/SpectrumX-Multiplanetech-blog-3840x2160-5205700-1-400x225.png 400w" sizes="auto, (max-width: 960px) 100vw, 960px" /></p>
<p><span style="font-weight: 400;">NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. </span><span style="font-weight: 400;">The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.</span></p>
<h3><b>Multiplane Unlocks Scale Without the Tradeoffs of a New Tier</b></h3>
<p><span style="font-weight: 400;">Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. </span><span style="font-weight: 400;">The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.</span></p>
<p><span style="font-weight: 400;">This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.</span></p>
<h3><b>Built Through Extreme Codesign</b></h3>
<p><span style="font-weight: 400;">That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives</span><span style="font-weight: 400;"> by 1.9x.</span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#scale-in"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="scale-in" class="wp-block-heading" style="scroll-margin-top: 100px;">NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA</h2>
<p><span style="font-weight: 400;">NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.</span></p>
<p><img loading="lazy" decoding="async" class="alignnone size-medium wp-image-97843" src="https://blogs.nvidia.com/wp-content/uploads/2026/08/bluefield-corp-blog-bluefield-4-1280x680-4468150-3-960x510.png" alt="" width="960" height="510" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/08/bluefield-corp-blog-bluefield-4-1280x680-4468150-3-960x510.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/08/bluefield-corp-blog-bluefield-4-1280x680-4468150-3-630x335.png 630w, https://blogs.nvidia.com/wp-content/uploads/2026/08/bluefield-corp-blog-bluefield-4-1280x680-4468150-3.png 1280w" sizes="auto, (max-width: 960px) 100vw, 960px" /></p>
<p><span style="font-weight: 400;">Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.</span></p>
<p><span style="font-weight: 400;">Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically. </span></p>
<p><span style="font-weight: 400;">Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign. </span></p>
<p><span style="font-weight: 400;">NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.</span></p>
<hr />
<p><em>Tuesday, Aug. 24, 8:00 a.m. PT <b><a href="https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#xpus"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f517.png" alt="🔗" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></b></em></p>
<h2 id="xpus" class="wp-block-heading" style="scroll-margin-top: 100px;">NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform</h2>
<p><span style="font-weight: 400;">NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.</span></p>
<p><span style="font-weight: 400;">As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.</span></p>
<p><span style="font-weight: 400;">The platform includes sixth-generation </span><a target="_blank" href="https://developer.nvidia.com/blog/nvidia-nvlink-the-scale-up-network-for-ai-factories/"><span style="font-weight: 400;">NVIDIA NVLink</span></a><span style="font-weight: 400;"> and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.</span></p>
<p><span style="font-weight: 400;">By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.</span></p>
<p><span style="font-weight: 400;">NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.</span></p>
# With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
Source: [https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/)
The next era of AI inference won’t be defined by a single breakthrough chip, network or system\. It’ll be defined by how every layer of the AI factory works together\. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems\.
[Announced today](https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai), the NVIDIA Vera Rubin rack\-scale system NVIDIA Groq 3 LPX is in full production\. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000\-token long\-context use cases critical to agentic systems, 4x faster than the nearest alternative platform\.
Industry partners worldwide are adopting Vera Rubin platform solutions\. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI\. CoreWeave has deployed into production Spectrum\-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high\-bandwidth, flat and lossless AI networks\. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX\.
As AI shifts from training to reasoning and agentic, inference has become the new frontier\. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems\.
These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale\.[https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/Rubin%20CPX%20Product%20Aug%202026.pdf](https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/Rubin%20CPX%20Product%20Aug%202026.pdf)
At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end\. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose\-built for the emerging demands of long\-context inference and multi\-agent systems\.
### **Extreme Codesign Optimizes for Performance**
Extreme codesign is the guiding principle behind NVIDIA platforms\. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences\.
[NVIDIA Spectrum\-X Ethernet](https://www.nvidia.com/en-us/networking/spectrumx/)moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds\. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture\.
NVIDIA Groq 3 LPX brings a new low\-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications\.
Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack\. From networking and context processing to large\-scale inference, NVIDIA’s full\-stack platform turns AI factories into integrated engines for intelligence, built to turn ever\-growing volumes of tokens into revenue\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#nvidia-partners)***
## NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs
Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications\.
Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real\-time AI experiences at scale\.
Connecting NVIDIA Vera Rubin racks,CoreWeaveis deploying Spectrum\-X Multiplane in production, unlocking advances for its AI cloud infrastructure\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#spacexai-vera-cpus)***
## **SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI**
SpaceXAIplans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites\. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU\-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation\.

The SpaceXAI[partnership](https://nvidianews.nvidia.com/news/spacexai-adopts-nvidia-vera-cpu-to-accelerate-agentic-ai-at-massive-scale)extends NVIDIA’s full\-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale\.
Designed for the agentic era, Vera Rubin provides leading per\-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#groq-3-lpx)***
## **NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator**
Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads\.
Agentic AI is creating a new performance challenge: decode latency\. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work\. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation\.
NVIDIA Rubin GPUs handle large\-scale context processing while LPX accelerates latency\-sensitive decode workloads\. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency\.

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large\-scale inference for the next generation of agentic AI applications\.
### **Building the Token Factory**
As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously\. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference\.
NVIDIA Groq 3 LPX was designed for exactly these workloads\. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization\.[https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/LPXDecks.pdf](https://nvidia-my.sharepoint.com/personal/scmartin_nvidia_com/Documents/Microsoft%20Copilot%20Chat%20Files/LPXDecks.pdf)
### **Extreme Codesign for Inference**
Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign\. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads\.
At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference\. A rack\-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip\-to\-chip links, creating a highly efficient inference engine built for modern AI factories\.
### **Designed for the Agentic AI Era**
As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve\. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose\-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories\.
And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming\. Stay tuned\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#spectrum-x-multiplane)***
## NVIDIA Spectrum\-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network
As AI factories grow massive, the network has become a critical engine of performance\. At Hot Chips, NVIDIA is spotlighting[Spectrum\-X Multiplane](https://www.nvidia.com/en-us/networking/spectrumx/)— the latest in the hardware\-accelerated Spectrum\-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier\.

NVIDIA Spectrum\-X Ethernet is designed as an end\-to\-end, AI\-optimized Ethernet platform, combining NVIDIA Spectrum\-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet\-based AI infrastructure for AI factories and clouds\.The platform is designed to deliver 1\.6x better AI networking performance compared with off\-the\-shelf Ethernet, while providing consistent, predictable performance in multi\-tenant environments\.
### **Multiplane Unlocks Scale Without the Tradeoffs of a New Tier**
Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power\. Spectrum\-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two\-tier network\.The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier\.
This all happens automatically\. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection\. In an eight\-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software\-based multiplane load balancing\. This translates to 1\.6x higher AI factory output\.
### **Built Through Extreme Codesign**
That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software\. Spectrum\-X SN6000 series switches, based on the 102\.4Tb/s Spectrum\-6 Ethernet ASIC and ConnectX\-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose\-built for Vera Rubin NVL72 AI factories\. Spectrum\-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super\-factory and accelerating multi\-site NCCL collectivesby 1\.9x\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#scale-in)***
## NVIDIA Introduces Scale\-In Infrastructure for Agentic AI Factories, Powered by BlueField\-4, DOCA
NVIDIA is introducing NVIDIA Scale\-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories\. Scale\-In extends purpose\-built acceleration to the infrastructure services that secure, manage and operate the AI factory\.

Powered by the NVIDIA BlueField\-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum\-X Ethernet, NVIDIA Scale\-In transforms the traditional north\-south access network into a unified, accelerated infrastructure domain\.
Cloud computing brought software\-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically\.
Agentic AI represents the next platform shift\. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services\. This transforms the demands on the infrastructure that brings AI to life\. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software\-defined flexibility with purpose\-built hardware acceleration and full\-stack codesign\.
NVIDIA Scale\-In delivers multi\-tenant networking, high\-performance storage access, in\-silicon security, elastic provisioning and real\-time observability, while keeping infrastructure processing independent of host compute resources\. By accelerating and codesigning these services as part of the AI factory, Scale\-In helps security, data access and operations scale alongside AI compute\. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale\.
---
*Tuesday, Aug\. 24, 8:00 a\.m\. PT**[🔗](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/#xpus)***
## NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform
NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world\-leading AI infrastructure platform, enabling hyperscalers and AI\-native companies to build semi\-custom AI factories with greater performance, flexibility and speed\.
As AI models grow in size and complexity, raw compute alone is not enough\. AI factories require high\-bandwidth, low\-latency scale\-up networking, proven rack\-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain\. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale\-up and scale\-out technology stack\.
The platform includes sixth\-generation[NVIDIA NVLink](https://developer.nvidia.com/blog/nvidia-nvlink-the-scale-up-network-for-ai-factories/)and NVLink Switch purpose\-built scale\-up networking, as well as NVLink\-C2C for energy\-efficient connectivity between XPUs and CPUs\. Through the NVIDIA MGX ecosystem, adopters can also use production\-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management\.
By standardizing GPU\- and XPU\-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness\. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve\.
NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon\. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone\.
NVIDIA has entered full production of Groq 3 LPX AI inference accelerator chips, which supercharge the Vera Rubin platform to achieve the fastest token generation speeds ever recorded for AI models.
NVIDIA is integrating Groq technology into rack-scale products to enhance agentic AI performance by splitting workloads across specialized processors, improving token generation latency and overall responsiveness.
NVIDIA announced its new Vera architecture CPUs, built from the ground up for agentic AI and reinforcement learning, claiming 2x performance over x86 alternatives. The Vera Rubin NVL72 platform integrates 72 GPUs and 36 CPUs, with major customers including Meta, Oracle, and Alibaba.
Groq raised $350 million at a $3.5 billion valuation after Nvidia licensed its technology and hired its senior team, pivoting to an inference cloud operation using both its LPUs and Nvidia systems.
NVIDIA's Vera Rubin NVL72 system demonstrates up to 30x higher throughput per megawatt for agentic AI workloads, setting a new efficiency standard for AI infrastructure.