Tag
This paper presents K2SVD, a principled method for learning Koopman operator representations with Kalman inference to achieve efficient and accurate time-series prediction, demonstrating superior performance over existing state-of-the-art approaches.
The paper introduces the 'Compression Trinity' framework, jointly applying sparsity, quantization, and low-rank approximations to compress Large Language Models for improved efficiency and performance.
This paper proposes a method for compressing large language models by combining neuron importance and data-aware low rank approximation, along with an efficient dynamic compression rate allocation algorithm, achieving performance on par with or better than previous state-of-the-art.