latent-attention

Tag

Cards List
#latent-attention

@shikhargupta02: I’ve been learning about latent attention (by deepseek). Instead of storing a full K and a V vector per token, it rathe…

X AI KOLs Timeline · 2026-08-08 Cached

The author shares insights from training a small model with DeepSeek's latent attention, observing layer-dependent latent usage and a test-time trick that reduces KV cache 4x without loss change.

0 favorites 0 likes
#latent-attention

Kimi K3 Architecture Overview and Notes

Hacker News Top · 2026-07-28 Cached

Sebastian Raschka provides an architectural overview of the open-weight Kimi K3 model, highlighting its scaling from 48B to 2.8T parameters, new LatentMoE and attention residual components, removal of RoPE in favor of NoPE, and native multimodal support. The model emphasizes inference efficiency and matches frontier performance.

0 favorites 0 likes
← Back to home

Submit Feedback