Tag
This paper presents a decomposition framework for quantization error in large language models, separating it into activation-guided weight compensation and orthogonal residual, and derives practical guidelines for improving W4A4 quantization through techniques like Hadamard rotation and sign selection.