@SemiAnalysis_: Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for …

X AI KOLs Following News

Summary

SemiAnalysis argues that Kimi K3's linear attention (KDA) is not detrimental to NVIDIA, HBM, DRAM, and networking, contrary to uninformed panic, and explains why reduced KV-cache requirements are actually beneficial.

Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 👇️ 1/8🧵 https://t.co/ih3RbtrDs8
Original Article
View Cached Full Text

Cached at: 07/20/26, 11:30 AM

Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 1/8

Kimi K3 is actually quite positive for NVIDIA, as large-model inference is where the NVL72 shines. Because K3 has more than 2.8 trillion parameters, it requires a large scale-up domain to store its weights. 2/8

Secondly, although Kimi Delta Attention has up to 10× lower networking requirements for KV-cache transfers, its large weights require even more network bandwidth to implement an optimization called WideEP, which spreads the weights across different GPUs. 3/8

WideEP distributes the 896 experts across many GPUs so that each GPU’s HBM contains only a small number of experts, optimizing per-token memory usage and compute utilization. 4/8

The unfortunate downside of the WideEP optimization is that it consumes a tremendous amount of network bandwidth. WideEP is highly optimized for rack-scale systems like the GB200/GB300 NVL72, whose copper backplane provides 18× more bandwidth than comparable DGX B200 systems. 5/8

Furthermore, since the weights occupy more than 1.5 TB of HBM capacity, the KV cache for K3’s KDA and Gated MLA will need to be offloaded to CPU DDR5 and NVMe, even at relatively low user concurrency, because little space remains in HBM. 6/8

Kimi themselves have stated that optimal K3 inferencing will require a rack witha n large scale up domain with at least 64 chips. 7/8

Lastly, Jevons’ Paradox means that making attention more efficient will drive wider AI adoption, which will ultimately require more GPUs, HBM, DRAM, and networking—not less. 8/8

Similar Articles

@elliotarledge: For those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is …

X AI KOLs Timeline

Elliot Arledge explains why he prefers using a Kimi Linear megakernel over Qwen 3.6 for kernel performance, comparing parameter counts, layer synchronization, hidden dimensions, and architecture-specific optimizations. The discussion highlights that Kimi Linear architecture is more suitable for megakernel implementation, especially for batch-1 decode on RTX PRO 6000 Blackwell.