@SemiAnalysis_: Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for …
Summary
SemiAnalysis argues that Kimi K3's linear attention (KDA) is not detrimental to NVIDIA, HBM, DRAM, and networking, contrary to uninformed panic, and explains why reduced KV-cache requirements are actually beneficial.
View Cached Full Text
Cached at: 07/20/26, 11:30 AM
Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 1/8
Kimi K3 is actually quite positive for NVIDIA, as large-model inference is where the NVL72 shines. Because K3 has more than 2.8 trillion parameters, it requires a large scale-up domain to store its weights. 2/8
Secondly, although Kimi Delta Attention has up to 10× lower networking requirements for KV-cache transfers, its large weights require even more network bandwidth to implement an optimization called WideEP, which spreads the weights across different GPUs. 3/8
WideEP distributes the 896 experts across many GPUs so that each GPU’s HBM contains only a small number of experts, optimizing per-token memory usage and compute utilization. 4/8
The unfortunate downside of the WideEP optimization is that it consumes a tremendous amount of network bandwidth. WideEP is highly optimized for rack-scale systems like the GB200/GB300 NVL72, whose copper backplane provides 18× more bandwidth than comparable DGX B200 systems. 5/8
Furthermore, since the weights occupy more than 1.5 TB of HBM capacity, the KV cache for K3’s KDA and Gated MLA will need to be offloaded to CPU DDR5 and NVMe, even at relatively low user concurrency, because little space remains in HBM. 6/8
Kimi themselves have stated that optimal K3 inferencing will require a rack witha n large scale up domain with at least 64 chips. 7/8
Lastly, Jevons’ Paradox means that making attention more efficient will drive wider AI adoption, which will ultimately require more GPUs, HBM, DRAM, and networking—not less. 8/8
Similar Articles
@elliotarledge: For those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is …
Elliot Arledge explains why he prefers using a Kimi Linear megakernel over Qwen 3.6 for kernel performance, comparing parameter counts, layer synchronization, hidden dimensions, and architecture-specific optimizations. The discussion highlights that Kimi Linear architecture is more suitable for megakernel implementation, especially for batch-1 decode on RTX PRO 6000 Blackwell.
@thealexker: underrated gems in Kimi-K3 release: > an early K3 wrote the majority of the kernels in the late development stages > it…
Kimi.ai released Kimi K3, a 2.8 trillion parameter multimodal model with 1 million context, featuring novel Delta Attention and Attention Residuals, and a self-optimizing stack including MiniTriton compiler. The model achieves up to 6.3x faster decoding and ~25% higher training efficiency.
@noisyb0y1: SOMEONE REVERSE-ENGINEERED KIMI K2.6 AND IT KILLS THE "BIGGER MODEL = BETTER AI" NARRATIVE FOR GOOD 1 trillion paramete…
A reverse engineering analysis of Kimi K2.6 reveals that its architecture prioritizes orchestration and skill injection over raw parameter count, achieving high SWE-Bench scores through multi-agent collaboration without retraining.
On Kimi K3: Its Capabilities And Related Discontents (70 minute read)
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.
@nrehiew_: > LatentMoE > 16 activated experts out of 896 > Kimi Delta Attention and AttnRes > 2.5x more efficient scaling This is …
Discussion of LatentMoE architecture with extreme sparsity (16/896 experts) and Kimi Delta Attention, claiming 2.5x more efficient scaling, and speculation about Kimi K3 model capabilities.