amd/Instella-MoE-16B-A3B-Think

Hugging Face Models Trending 模型

摘要

AMD发布Instella-MoE,一个完全开源的160亿参数混合专家语言模型,其中28亿参数被激活,从头使用AMD Instinct GPU训练,并发布了从预训练到强化学习的所有训练阶段。

任务:文本生成 标签:transformers, safetensors, deepseek_v3, 文本生成, instella, moe, amd, rocm, 对话, custom_code, en, arxiv:2511.10628, license:other, 文本生成推理, endpoints_compatible, region:us
查看原文
查看缓存全文

缓存时间: 2026/07/30 03:54

amd/Instella-MoE-16B-A3B-Think · Hugging Face

来源:https://huggingface.co/amd/Instella-MoE-16B-A3B-Think

Instella-MoE✨:完全开源的最先进混合专家语言模型

Instella-MoE 是一个完全开源的最先进混合专家(MoE)语言模型,总参数量为 160 亿,每个 token 激活参数量为 28 亿,从预训练到强化学习(RL)实现了端到端训练。该模型从零开始在 AMD Instinct™ MI300X 和 MI325X GPU 上,使用 AMD 的 Primus 框架训练而成。Instella-MoE 将稀疏激活的 MoE 设计与架构创新相结合,例如门控多头潜在注意力(Gated MLA)和 FarSkip-Collective(https://github.com/AMD-AGI/FarSkip-Collective)。

Instella-MoE 成本与性能对比

图 1:预训练和后训练的 Instella-MoE 模型性能与其他类似大小的最先进模型对比。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#takeaways 要点

  • Instella-MoE 是 AMD 开发的新型最先进完全开源混合专家语言模型,总参数量为 160 亿,每个 token 激活参数量为 28 亿,从零开始在 AMD Instinct™ MI300X 和 MI325X GPU 上训练而成。
  • Instella-MoE 模型检查点发布涵盖了模型训练管道的各个阶段,包括预训练、中期训练、长上下文扩展、SFT、DPO 和 RL。
  • Instella-MoE 完全构建在 AMD ROCm™ 软件栈之上,基于 Primus 训练框架和 Miles RL 框架,并融入了最先进的架构和系统创新——包括门控多头潜在注意力(Gated MLA)以及通过 FarSkip-Collective 实现的极致通信-计算重叠——从而在 AMD 硬件上实现高效的大规模训练和推理。
  • 完全开源且可访问:我们提供了所有训练阶段的完整训练配方,包括训练框架、数据混合、中间检查点和推理代码。

本次发布的检查点来自 Instella-MoE 训练管道的以下阶段,如下表 1 所示:

模型阶段描述
Instella-MoE-16B-A3B-Pretrain(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-Pretrain)预训练从零开始在大型多样化训练语料库上训练的 MoE 基础模型。
Instella-MoE-16B-A3B-Midtrain(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-Midtrain)中期训练在高质量数据混合上进一步训练预训练模型,以优化关键能力。
Instella-MoE-16B-A3B-Base(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-Base)长上下文长上下文训练以扩展模型处理和推理更长序列的能力。我们将此作为最终的基础检查点
Instella-MoE-16B-A3B-SFT(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-SFT)SFT通过监督微调(SFT)扩展基础检查点,以启用指令跟随和思维链推理能力。
Instella-MoE-16B-A3B-DPO(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-DPO)DPO在对比偏好数据上进行直接偏好优化(DPO)以提升模型性能。
Instella-MoE-16B-A3B-Think(链接)(https://huggingface.co/amd/Instella-MoE-16B-A3B-Think)RL使用强化学习(RL)优化的最终思考检查点,进一步强化指令跟随能力和整体响应质量。

***表 1:*Instella-MoE-16B-A3B 模型及训练阶段。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#model-summary 模型摘要

参数
总参数量16B
每个 Token 激活参数量2.8B
解码器层数27
隐藏层大小2048
注意力头数16
专家数量64
共享专家数2
每个 Token 激活专家数6
词表大小128,896
注意力机制门控多头潜在注意力(Gated MLA)
MoE 连接方式FarSkip-Collective

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#results 结果

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#pretraining-results 预训练结果

基础模型在标准基准上的结果

表 2:Instella-MoE-16B-A3B-Base 在标准基准上的结果。

基础模型在长上下文基准上的结果

表 3:Instella-MoE-16B-A3B-Base 在长上下文 HELMET 和 RULER 基准上的结果。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#post-training-results 后训练结果

Instella-MoE-Think 后训练结果

表 4:Instella-MoE-Think 结果。我们使用 OLMES 框架评估所有模型,最多生成 32768 个 token。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#getting-started 开始使用

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#example-usage 使用示例

from transformers import AutoModelForCausalLM, AutoTokenizer
checkpoint = "amd/Instella-MoE-16B-A3B-Think"

tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(checkpoint, device_map="auto", trust_remote_code=True)

prompt = [{"role": "user", "content": "What are the computational benefits of Mixture-of-Experts models?"}]
inputs = tokenizer.apply_chat_template(
    prompt,
    add_generation_prompt=True,
    return_tensors='pt'
)

tokens = model.generate(
    inputs.to(model.device),
    max_new_tokens=1024,
    temperature=0.6,
    top_p=0.95,
    do_sample=True
)

print(tokenizer.decode(tokens[0], skip_special_tokens=False))

如需使用 SGLang(https://github.com/sgl-project/sglang)进行高吞吐量推理,请参阅我们 GitHub 仓库(https://github.com/AMD-AGI/Instella-MoE)中的设置和使用说明。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#training-details 训练细节

Instella-MoE 在 AMD Instinct™ MI300X 和 MI325X GPU 上,使用 AMD ROCm™ 软件栈,基于 Primus(https://github.com/AMD-AGI/Primus)训练框架和 Miles(https://github.com/radixark/miles)RL 框架进行端到端训练。训练过程分为多阶段流水线——预训练、中期训练、长上下文扩展、SFT、DPO 和 RL——每个阶段逐步增强模型的能力。

关于完整的训练配方,包括各阶段数据混合、超参数、训练框架和推理代码,请参阅我们的 GitHub 仓库(https://github.com/AMD-AGI/Instella-MoE)和技术博客(https://rocm.blogs.amd.com/artificial-intelligence/instella-moe/README.html)。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#acknowledgements 致谢

我们衷心感谢 LLM360 团队和 Miles 团队在模型开发过程中给予的宝贵支持。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#license 许可

  • Instella-MoE 模型根据 ResearchRAIL 许可在学术和研究目的下授权使用。
  • 有关更多信息,请参阅 LICENSE(https://huggingface.co/amd/Instella-MoE-16B-A3B-Think/tree/main/LICENSE)。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#bias-risks-and-limitations 偏见、风险与限制

  • 这些模型仅用于研究目的。它们不适用于需要高事实准确性、安全关键应用或健康医疗应用的使用场景。不得用于生成虚假信息或促进有毒对话。
  • 模型检查点是在没有任何安全承诺的情况下提供的。用户务必根据各自的使用场景进行全面评估并实施安全过滤机制。
  • 可能会诱导模型生成事实不准确、有害、暴力、有毒、带有偏见或其他令人反感的内容。这些内容也可能出现在并非旨在引发此类响应的提示中。因此,请用户注意这一点,并在使用模型时谨慎行事并秉持负责任的态度。
  • 模型的多语言能力尚未经过测试,因此可能在不同语言中误解并产生错误响应。

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#contributors 贡献者

核心贡献者:Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Zicheng Liu
贡献者:Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Emad Barsoum

https://huggingface.co/amd/Instella-MoE-16B-A3B-Think#citations 引用

Instella-MoE 技术报告即将发布。在此期间,欢迎引用:

@article{instella,
  title={Instella: Fully Open Language Models with Stellar Performance},
  author={Liu, Jiang and Wu, Jialian and Yu, Xiaodong and Su, Yusheng and Mishra, Prakamya and Ramesh, Gowtham and Ranjan, Sudhanshu and Manem, Chaitanya and Sun, Ximeng and Wang, Ze and Brahma, Pratik Prabhanjan and Liu, Zicheng and Barsoum, Emad},
  journal={arXiv preprint arXiv:2511.10628},
  year={2025}
}

@inproceedings{
dukler2026farskipcollective,
title={FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models},
author={Yonatan Dukler and Guihong Li and Deval Shah and Jiang Liu and Vikram Appia and Emad Barsoum},
booktitle={Ninth Conference on Machine Learning and Systems},
year={2026},
url={https://openreview.net/forum?id=ruOpvLzsGV}
}

相似文章

AMD Instella-MoE-16B-A3B

Reddit r/LocalLLaMA

AMD 在 HuggingFace 上发布了一个开源的混合专家模型 Instella-MoE-16B-A3B。

AI2推出的新MoE模型:EMO

Reddit r/LocalLLaMA

AI2发布了EMO,一个混合专家(MoE)语言模型,总参数量14B,其中1B活跃参数,基于1万亿tokens训练,并采用文档级路由,即专家会按领域(如健康、新闻等)进行聚类。

MobileMoE:扩展端侧混合专家模型

Hugging Face Daily Papers

MobileMoE 引入了高效的端侧混合专家语言模型,参数规模低于十亿,在性能和效率上均优于密集基线模型和现有的 MoE 模型。这些模型在开源数据集上训练,并在商用智能手机上展现出显著的加速效果。