Tag
NVIDIA's Vera Rubin NVL72 system debuts with leading performance in MLPerf Inference v6.1, delivering up to 3.7x better throughput than GB300 NVL72 on benchmarks like DeepSeek-R1 and Qwen3-VL, emphasizing system performance, scaling efficiency, and software optimizations.
Cursor is open-sourcing Mixture-of-Kittens (MoK), a production MoE training megakernel for NVL72s that fuses communication and computation, delivering a 1.41x end-to-end training throughput improvement for their Composer model.
Cursor is open-sourcing Mixture-of-Kittens (MoK), a fused MoE training megakernel for NVIDIA NVL72 that runs up to 2.37x faster than the strongest public baselines.
SpaceX raised the peak power spec of its AI Satellite V1 to ~250kW (battery-assisted) with ~160kW average, enabling it to handle an NVL72 Ruben rack.
FastAFD is an open-source serving system for Attention-FFN Disaggregation of MoE models on Blackwell NVL72, achieving 1.35-1.45× per-GPU decode throughput improvement over colocated MoE serving.