neural-network-compression

Tag

Cards List
#neural-network-compression

Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression

arXiv cs.LG · 2026-07-22 Cached

This paper proposes a method for compressing large language models by combining neuron importance and data-aware low rank approximation, along with an efficient dynamic compression rate allocation algorithm, achieving performance on par with or better than previous state-of-the-art.

0 favorites 0 likes
#neural-network-compression

Empirical Minimal-Realisation Compression of Deep Neural Networks via Controllability-Observability Tests

arXiv cs.LG · 2026-07-08 Cached

This paper proposes a controllability–observability framework for compressing deep neural networks by reducing hidden-state redundancy, demonstrating significant compression with minimal accuracy loss on MNIST and CIFAR-10.

0 favorites 0 likes
#neural-network-compression

Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) for Exponential Compression of Deep Neural Networks

arXiv cs.LG · 2026-06-02 Cached

Introduces Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) for compressing deep neural network layers via small core tensors, achieving high compression ratios while maintaining accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback