Tag
This paper proposes a method for compressing large language models by combining neuron importance and data-aware low rank approximation, along with an efficient dynamic compression rate allocation algorithm, achieving performance on par with or better than previous state-of-the-art.
This paper proposes a controllability–observability framework for compressing deep neural networks by reducing hidden-state redundancy, demonstrating significant compression with minimal accuracy loss on MNIST and CIFAR-10.
Introduces Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) for compressing deep neural network layers via small core tensors, achieving high compression ratios while maintaining accuracy.