Tag
NVIDIA is integrating Groq technology into rack-scale products to enhance agentic AI performance by splitting workloads across specialized processors, improving token generation latency and overall responsiveness.
Cerebras launches the CS-4, a rack-scale AI system with WSE-3 Turbo technology claiming up to 30x faster inference than GPUs, featuring a modular design for efficient hyperscale deployment.
AMD launches Helios, its first rack-scale AI system to rival Nvidia's offerings, with Microsoft joining as a customer. The system is set to ship later this year.
NVIDIA argues that performance per watt is the key metric for AI infrastructure efficiency, highlighting how its Blackwell and Vera Rubin platforms achieve up to 25x improvement over Hopper for MoE models.