low-rank-projection

Tag

Cards List
#low-rank-projection

@shikhargupta02: I’ve been learning about latent attention (by deepseek). Instead of storing a full K and a V vector per token, it rathe…

X AI KOLs Timeline ↗ · 2026-08-08 Cached

The author shares insights from training a small model with DeepSeek's latent attention, observing layer-dependent latent usage and a test-time trick that reduces KV cache 4x without loss change.

0 favorites 0 likes
#low-rank-projection

OjaKV: Context-Aware Online Low-Rank KV Cache Compression

arXiv cs.CL ↗ · 2026-04-20 Cached

OjaKV introduces a context-aware online low-rank KV cache compression framework that uses hybrid storage and Oja's algorithm for incremental subspace adaptation to reduce GPU memory bottlenecks in long-context LLM inference without model fine-tuning.

0 favorites 0 likes
← Back to home

Submit Feedback