NVIDIA reports up to 30× more agentic throughput per MW on Vera Rubin—but tokens/MW still is not completed work/MW
Summary
NVIDIA reports up to 30× higher agentic throughput per megawatt with Vera Rubin compared to GB300, but emphasizes that production benchmarks should include additional metrics beyond tokens per watt.
Similar Articles
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
NVIDIA's Vera Rubin NVL72 system demonstrates up to 30x higher throughput per megawatt for agentic AI workloads, setting a new efficiency standard for AI infrastructure.
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
NVIDIA announces Vera Rubin platform, delivering 10x more throughput per megawatt than Blackwell and claiming lowest token cost for AI factories through extreme co-design across seven chips and five rack trays.
NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads — a Key Metric for Agentic AI
NVIDIA introduces Vera Rubin, a new architecture designed to maximize intelligence per dollar for post-training workloads, addressing the continuous learning demands of agentic AI by optimizing cost per token and supporting reinforcement learning at scale.
NVIDIA Enters Full Production of Groq 3 LPX AI Inference Accelerator Chips, Supercharging Vera Rubin With The Fastest Token Generation Speeds Ever Recorded (4 minute read)
NVIDIA has entered full production of Groq 3 LPX AI inference accelerator chips, which supercharge the Vera Rubin platform to achieve the fastest token generation speeds ever recorded for AI models.
NVIDIA Next Gen Vera Rubin GPUs scheduled for mid-2027
NVIDIA expects $20 billion in sales from Vera Rubin systems in Q3 FY2027, marking its fastest product ramp in history, with shipments beginning to major hyperscalers like Microsoft.