linear-layers

Tag

Cards List
#linear-layers

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

arXiv cs.LG ↗ · 5d ago Cached

PRQuant is a training-free and low-overhead framework for quantizing linear layers in large language models, using permutation and residual compensation to reduce inference latency while improving accuracy over baselines like MXFP4.

0 favorites 0 likes
← Back to home

Submit Feedback