long-context-inference

Tag

Cards List
#long-context-inference

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

arXiv cs.LG · yesterday Cached

PuzzleKV is a training-free method for compressing key-value cache in large language models using page-wise low-rank decomposition, achieving over 96% performance with approximately 60% storage.

0 favorites 0 likes
← Back to home

Submit Feedback