Tag
The paper identifies activation quantization as the primary bottleneck in ultra-low-bit quantized multimodal LLMs and proposes ResidualFallbackQuantization (RFQ) to recover performance with minimal overhead.