low-rank-decomposition

Tag

Cards List
#low-rank-decomposition

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

arXiv cs.LG · 23h ago Cached

PuzzleKV is a training-free method for compressing key-value cache in large language models using page-wise low-rank decomposition, achieving over 96% performance with approximately 60% storage.

0 favorites 0 likes
#low-rank-decomposition

Break Through the Compression Bottleneck: From Theory to Practice

arXiv cs.CL · 2026-07-24 Cached

This paper provides the first mathematical proof that low-rank decomposition and quantization are non-orthogonal when combined for LLM compression, leading to performance degradation, and proposes a novel Diagonal Adhesive Method (DAM) to mitigate this loss.

0 favorites 0 likes
#low-rank-decomposition

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

arXiv cs.CL · 2026-06-08 Cached

Introduces SigmaScale, a method that learns auxiliary scaling matrices for SVD-based LLM compression, showing competitive performance on Llama 3.1 8B and Qwen3-8B benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback