bit-allocation

标签

Cards List
#bit-allocation

Colla-Q:通过最小最大精度平衡在MoE量化中实现协作专家

arXiv cs.LG ↗ · 2026-09-17 缓存

本文介绍了Colla-Q,一种基于激活熵的比特分配框架,用于量化混合专家模型,以平衡专家性能,提高整体模型效率并减少对校准数据集的依赖。

0 人收藏 0 人点赞
#bit-allocation

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

Reddit r/LocalLLaMA ↗ · 2026-08-09 缓存

KLQ is a training-free LLM quantization method that allocates bits per direction based on measured KL divergence, outperforming existing training-free rotation-based methods on W4A4KV4-bit settings for models like Llama 3.2 1B and Qwen 2.5.

0 人收藏 0 人点赞
#bit-allocation

感知RoPE的KV缓存量化比特分配方法

arXiv cs.LG ↗ · 2026-06-24 缓存

提出Block-GTQ,一种感知RoPE的KV缓存量化比特分配方法,通过为高能量RoPE块分配更多比特,提升长上下文性能与内存效率。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈