qkv-projection

Tag

Cards List
#qkv-projection

Do transformers need three projections? Systematic study of QKV variants

Hacker News Top · 2026-06-04 Cached

This paper systematically studies variants of QKV projection sharing in transformers, finding that sharing key and value projections (Q-K=V) achieves 50% KV cache reduction with only 3.1% perplexity degradation, and combining with GQA/MQA can reach up to 96.9% cache reduction—enabling practical on-device inference with minimal quality loss.

0 favorites 0 likes
← Back to home

Submit Feedback