NVIDIA reports up to 30× more agentic throughput per MW on Vera Rubin—but tokens/MW still is not completed work/MW

Reddit r/ArtificialInteligence News

Summary

NVIDIA reports up to 30× higher agentic throughput per megawatt with Vera Rubin compared to GB300, but emphasizes that production benchmarks should include additional metrics beyond tokens per watt.

NVIDIA’s new AgentX results replay production-style coding-agent sessions with long context, KV-cache reuse, tool gaps, and dynamic concurrency. Its Vera Rubin preview result claims up to 30× higher throughput per megawatt than GB300 at a matched interactivity target; NVIDIA says the result is pending SemiAnalysis review. This is more representative than fixed 8K/1K serving, but the numerator still stops at tokens. A production benchmark should also report: - Accepted task outcomes per MWh - Completion-latency distribution - Tool and retry amplification - Cache hit rate and memory pressure - Model and harness equivalence - Human review minutes and rollback rate An efficient system can generate more unusable work just as efficiently. Source: NVIDIA Technical Blog, August 24, 2026 — https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/
Original Article

Similar Articles