Tag
SemiAnalysis argues that Kimi K3's linear attention (KDA) is not detrimental to NVIDIA, HBM, DRAM, and networking, contrary to uninformed panic, and explains why reduced KV-cache requirements are actually beneficial.
Enze Xie announces Sol Video Inference Engine, an agent-native, training-free full-stack accelerator for video diffusion that auto-tunes cache, sparse attention, token pruning, quantization, and kernel fusion, achieving >2× end-to-end speedup on large models like 64B Cosmos3-Super and 22B LTX-2.3.