标签
本文提出校准裁剪方法,通过将FP8裁剪界限与高精度分布对齐,来稳定LLMs强化学习中的FP8量化,消除熵激增并恢复性能。
dealignai released a cybersecurity-focused weight-modified variant of the 753B-parameter GLM-5.3-FP8 model that reduces refusals for offensive-security tasks such as red-teaming, exploit development, and malware analysis while preserving base capabilities and running natively on Hopper GPUs via stock vLLM.
本文发布了Qwen3.8-Flash-Next模型的FP8量化权重,介绍了如混合注意力机制(Hybrid Attention with QSA)、门控残差(Gated Residual)和N-gram嵌入(N-gram Embedding)等架构创新,以提高大型语言模型的效率。
本文介绍了一个未审查的Qwen3.8 27B AI模型,该模型经过修改以移除安全限制,并量化为FP8以提高效率,适用于研究目的。
这是Qwen3.8-27B的修改版本,移除了安全拒绝机制并进行了FP8量化,专为AI安全性和可解释性研究设计。
Qwen releases FP8-quantized weights for Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model with 95B activated parameters, claiming Qwen-Max-class capability in an open release with strong coding, agentic, and long-context performance.
量化版 27B Qwen3.6 在单颗 49W GB10 GPU 上借助 Dflash+DDTree 优化,256k 上下文、10 智能体并发,峰值达 200 tok/s,平均 136 tok/s。
阿里巴巴发布 Qwen3.6-27B-FP8,一款 27B 参数的 FP8 量化模型,在代理式编码与推理基准上表现强劲,现已上架 Hugging Face。