language-model-compression

Tag

Cards List
#language-model-compression

Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression

arXiv cs.LG · 2d ago Cached

This paper proposes a method for compressing large language models by combining neuron importance and data-aware low rank approximation, along with an efficient dynamic compression rate allocation algorithm, achieving performance on par with or better than previous state-of-the-art.

0 favorites 0 likes
← Back to home

Submit Feedback