@shikhargupta02: I’ve been learning about latent attention (by deepseek). Instead of storing a full K and a V vector per token, it rathe…
Summary
The author shares insights from training a small model with DeepSeek's latent attention, observing layer-dependent latent usage and a test-time trick that reduces KV cache 4x without loss change.
View Cached Full Text
Cached at: 08/08/26, 01:04 PM
I’ve been learning about latent attention (by deepseek). Instead of storing a full K and a V vector per token, it rather stores a common latent vector L for both - using a learned low rank projection of the input. K and V can be constructed from L by projecting it up (again, learned). The dimension of this latent vector << total K+V storage per token. So naturally, the stored KV cache goes down.
Some interesting patterns emerged when I trained a small model with it. The first and the final few layers rely on the latent space a lot more than the intermediate layers. Layers 5-7 use only 25% of the 256 dimensions available. They seem to need less complexity in their K and V vectors. The first layer, on the other hand, uses the latent space the most which makes sense as it is closest to the embedding layer and has a lot of raw signal to encode.
I did a test time experiment where I keep only the most prominent latent dims and shave off the rest. There was almost no change in the validation loss while reducing the kv cache 4x compared to full rank latent vector (which was already 6x less compared to multi head attention). A very cool test time optimization.
Similar Articles
@omarsar0: Banger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated to…
This paper introduces Declarative Attention, a protocol that allows language models to declare where to attend in their chain-of-thought, reducing attended tokens by 52% on Gemma-4-31B and 31.1% on Qwen-3.6-27B with minimal accuracy drops.
@thtrkim: Visual deep dive on FlashAttention by hand (drawn with Excalidraw) https://winterrykim.github.io/blog/2026/training-lm-…
A visual deep dive into FlashAttention, explaining memory optimization and operator fusion for efficient attention computation in language model training.
@modal: DeepSeek-V4-Flash has 284B total parameters with 13B active per token. Combined with a hybrid compressed attention mech…
DeepSeek-V4-Flash is a 284B-parameter MoE model with 13B active parameters per token, featuring a hybrid compressed attention mechanism that reduces KV cache needs for 1M-token contexts. It can be served with SGLang on Modal for fast decoding on a single B300.
@omarsar0: NEW paper worth reading. (bookmark it) The basic idea is to pair a compressive recurrent state with a small exact memor…
HOLA (Hippocampal Linear Attention) augments linear attention with a bounded exact KV cache inspired by hippocampal memory, improving long-range recall and perplexity without sacrificing efficiency. At 340M parameters, it outperforms full-attention Transformers on Wikitext and achieves robust needle recall up to 32k tokens.
@Michaelzsguo: KV cache is the model’s working memory during generation. As the context window gets longer, the model has to keep more…
DeepSeek's KV cache compression innovations, including MLA and CSA/HCA, reduce KV cache size by 93%, enabling efficient long-context inference and SSD-based caching, as demonstrated by antirez's ds4.c project.