openbmb/MiniCPM5-2B

Hugging Face Models Trending 模型

摘要

OpenBMB发布了MiniCPM5-2B,这是一个2B参数的密集Transformer模型,在设备端部署的同类模型中达到了最先进的性能,并提供了开源的高质量训练数据集。

任务:文本生成 标签:transformers, safetensors, llama, 文本生成, minicpm, minicpm5, 长上下文, 工具调用, 设备端, 边缘AI, 对话式, en, zh, 数据集:openbmb/Ultra-FineWeb, 数据集:openbmb/UltraX-Preview, 数据集:openbmb/Ultra-FineWeb-L3, 数据集:openbmb/UltraData-Math, 数据集:openbmb/UltraData-Code, 数据集:openbmb/UltraData-SFT-2605, 数据集:openbmb/UltraData-SFT-Agent-2609, 数据集:openbmb/UltraData-RL-2609, arxiv:2506.07900, arxiv:2602.09003, 许可:apache-2.0, 文本生成推理, 端点兼容, 地区:us
查看原文
查看缓存全文

缓存时间: 2026/09/08 00:07

openbmb/MiniCPM5-2B · Hugging Face

来源:https://huggingface.co/openbmb/MiniCPM5-2B

MiniCPM技术报告(https://arxiv.org/pdf/2506.07900)|MiniCPM维基(中文)(https://modelbest.feishu.cn/wiki/UtWxwcERfiRIpIkBOjuc3h9tn1D)|GitHub仓库(https://github.com/OpenBMB/MiniCPM)|UltraData(https://ultradata.openbmb.cn/)|在线演示(https://huggingface.co/spaces/openbmb/MiniCPM5-2B-Demo)

English |中文(https://huggingface.co/openbmb/MiniCPM5-2B/blob/main/README-cn.md)

https://huggingface.co/openbmb/MiniCPM5-2B#highlights亮点

我们发布了MiniCPM5-2B,这是MiniCPM5系列的第二款模型,继MiniCPM5-1B(https://huggingface.co/openbmb/MiniCPM5-1B)之后。它是一个基于Transformer的稠密2B参数模型,采用了相同的训练方案进行扩展,专为本地部署、端侧使用和资源受限场景设计,达到了2B参数级别开源模型的SOTA水平。

🏆2B级别开源SOTA:与同等规模的强劲开源模型相比,MiniCPM5-2B在该比较集中达到了SOTA性能。整体上它仍与4B级别的模型具有竞争力,同时在编程、数学、长上下文理解、工具使用和智能体任务方面展现出超越同规模模型的优势。

📂开放高质量数据:伴随模型发布,我们还将背后的高质量训练数据集作为UltraData(https://ultradata.openbmb.cn/)家族的一部分开源:UltraX(https://huggingface.co/datasets/openbmb/UltraX-Preview),一个高质量的网络预训练数据集;UltraData-Code(https://huggingface.co/datasets/openbmb/UltraData-Code),具有L0-L3层级代码数据管理,显著提升编码能力;UltraData-SFT-Agent-2609(https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609),包含50万条智能体训练样本,增强全面的本地智能体能力;以及UltraData-RL-2609(https://huggingface.co/datasets/openbmb/UltraData-RL-2609),拥有超过8万条高质量RL训练样本,覆盖数学、代码、通用知识和长上下文推理。

https://huggingface.co/openbmb/MiniCPM5-2B#model-list模型列表

使用此目录选择与您运行时环境匹配的模型格式:

MiniCPM5-2B

  • MiniCPM5-2B(https://huggingface.co/openbmb/MiniCPM5-2B)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B)· BF16最终发布版(使用RL + OPD后训练)👈 您在此处
  • MiniCPM5-2B-SFT(https://huggingface.co/openbmb/MiniCPM5-2B-SFT)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-SFT)· BF16纯SFT检查点(RL / OPD之前)
  • MiniCPM5-2B-Midtrain(https://huggingface.co/openbmb/MiniCPM5-2B-Midtrain)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-Midtrain)· BF16中期训练检查点(SFT之前)
  • MiniCPM5-2B-Base(https://huggingface.co/openbmb/MiniCPM5-2B-Base)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-Base)· BF16基础检查点(仅预训练)
  • MiniCPM5-2B-GGUF(https://huggingface.co/openbmb/MiniCPM5-2B-GGUF)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-GGUF)· 适用于llama.cpp / Ollama / LM Studio的GGUF格式
  • MiniCPM5-2B-MLX(https://huggingface.co/openbmb/MiniCPM5-2B-MLX)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-MLX)· 适用于Apple Silicon的MLX / 4bit格式
  • MiniCPM5-2B-GPTQ(https://huggingface.co/openbmb/MiniCPM5-2B-GPTQ)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-GPTQ)· GPTQ / 4bit量化模型
  • MiniCPM5-2B-DSpark(https://huggingface.co/openbmb/MiniCPM5-2B-DSpark)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-DSpark)· 用于推理加速的DSpark草稿模型

MiniCPM5-1B

  • MiniCPM5-1B(https://huggingface.co/openbmb/MiniCPM5-1B)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B)· BF16最终发布版(使用RL + OPD后训练)
  • MiniCPM5-1B-SFT(https://huggingface.co/openbmb/MiniCPM5-1B-SFT)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-SFT)· BF16纯SFT检查点(RL / OPD之前)
  • MiniCPM5-1B-Base(https://huggingface.co/openbmb/MiniCPM5-1B-Base)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-Base)· BF16基础检查点(仅预训练)
  • MiniCPM5-1B-GGUF(https://huggingface.co/openbmb/MiniCPM5-1B-GGUF)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-GGUF)· 适用于llama.cpp / Ollama / LM Studio的GGUF格式
  • MiniCPM5-1B-MLX(https://huggingface.co/openbmb/MiniCPM5-1B-MLX)·ModelScope(https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-MLX)· 适用于Apple Silicon的MLX / 4bit格式

https://huggingface.co/openbmb/MiniCPM5-2B#model-information模型信息

MiniCPM5-2B具有以下特性:

  • 类型:因果语言模型
  • 架构:标准的LlamaForCausalLM
  • 参数数量:2,516,756,480
  • 非嵌入参数数量:1,981,982,720
  • 层数:42
  • 注意力头数(GQA):Q头16个,KV头2个
  • 上下文长度:131,072

https://huggingface.co/openbmb/MiniCPM5-2B#introduction简介

MiniCPM5-2B是MiniCPM5系列中的第二款模型。它专为本地助手、编码智能体、工具使用工作流以及需要紧凑模型的推理场景而设计。该模型在保持较小部署占用的同时,提供了原生的长上下文支持。

https://huggingface.co/openbmb/MiniCPM5-2B#evaluation-results评测结果

我们将MiniCPM5-2B与同尺寸类别中的强劲开源模型进行了比较,包括LFM2.5-2.6BQwen3.5-2BGemma-4-E2B-it,同时列出了更大的模型如Qwen3.5-4Bgranite-4.2-3BNemotron-3-Nano-4BGemma-4-E4B-itLFM2.5-8B-A1B作为参考。

在该比较集中,MiniCPM5-2B以平均得分53.9达到了2B级别开源SOTA,并且超过了此处列出的所有更大模型(最高分为51.1)。其优势在代码推理、数学推理、长上下文理解、工具使用和多项智能体任务中最为明显。

MiniCPM5-2B与基线模型评测结果

MiniCPM5-2B2B级别模型4B级别模型LFM2.5-2.6BQwen3.5-2BGemma-4-E2B-itQwen3.5-4Bgranite-4.2-3BNemotron-3-Nano-4BGemma-4-E4B-itLFM2.5-8B-A1B平均分

53.933.228.024.651.142.732.631.228.4代码推理LiveCodeBench v6

69.142.120.242.956.458.950.753.939.8LCB-Pro 25Q2(简单)

68.030.910.327.158.354.651.645.827.8LCB-Pro 25Q2(中等)

17.50.00.00.07.05.35.31.80.0OJBench

32.511.22.611.624.821.820.019.08.2SciCode(wbg)

26.3†14.2†2.8†20.9†16.1†24.9†16.4†24.4†7.8†数学推理AIME 2025

86.541.929.631.778.879.456.337.146.0AIME 2026

86.545.229.039.882.783.562.145.056.7HMMT 2026年2月

63.833.720.517.864.060.851.330.138.5MATH-500

94.689.685.885.499.097.091.688.293.2指令遵循IFBench

66.359.046.025.759.073.058.328.351.0IFEval

86.793.477.531.490.293.788.044.490.8Multi-IF

71.876.857.140.373.675.965.945.971.4通用知识MMLU-Pro

70.865.264.356.078.065.865.768.363.1MMLU-Redux

84.780.080.071.888.778.979.883.780.0HLE

8.9†6.2†2.6†4.8†9.9†6.6†4.9†3.8†6.9†GPQA-Diamond

70.2†55.8†45.6†43.3†77.1†55.9†51.3†57.6†51.3†SuperGPQA

40.826.238.630.352.839.937.838.734.5长上下文AA-LCR

59.0†5.3†28.7†17.0†61.0†24.3†17.3†33.0†0.0†NoLiMa

68.10.717.13.943.55.11.12.30.5LongBenchPro

44.823.78.242.258.434.827.953.519.6LongBench v2

43.730.324.933.247.336.032.042.730.4工具使用τ3-Bench Banking

20.8†7.2†2.13.96.8†5.6†1.24.13.4τ2-Bench Telecom

97.190.469.0†20.8†92.1†40.928.1†20.8†16.1†BFCL v4

66.661.143.636.656.852.243.747.049.2编码智能体SWE-bench Verified

46.46.05.02.033.636.83.015.00.4SWE-bench Pro

14.40.60.80.028.212.30.13.30.4Terminal-Bench v2.1

8.6†4.5†3.0†0.4†25.8†13.9†3.8†1.9†1.9搜索智能体BrowseComp-ZH

43.59.818.24.739.621.13.37.013.2BrowseComp Top100

39.713.719.36.033.319.04.76.39.7GAIA Text-103

88.749.547.930.178.657.326.539.541.1通用智能体GDPval-AA v2

19.6†4.50.00.011.70.0†0.00.00.0Claw-Gym

59.219.325.531.351.660.033.737.92.7WildClaw

23.910.29.28.917.020.08.914.34.5QwenClaw

42.919.318.214.537.136.416.816.74.51.蓝色粗体表示该行所有模型(包括4B级别模型)中的最佳结果;黑色粗体表示2B级别模型中的最佳结果。2. 标有†的分数来自官方Artificial Analysis发布;其他分数均为内部复现。

https://huggingface.co/openbmb/MiniCPM5-2B#training-recipe训练方案

MiniCPM5-2B的训练是**UltraData分层数据管理(https://arxiv.org/pdf/2602.09003)**的完整实践,涵盖三个阶段:基础训练、中期训练和后训练。

基础训练阶段,模型经过稳定训练和衰减训练以构建核心语言能力和训练稳定性。然后进入中期训练,进一步强化目标能力并适应目标数据分布。训练语料库与模型一同发布,包括Ultra-FineWeb(https://huggingface.co/datasets/openbmb/Ultra-FineWeb)、Ultra-FineWeb-L3(https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3)、UltraX(https://huggingface.co/datasets/openbmb/UltraX-Preview)、UltraData-Code(https://huggingface.co/datasets/openbmb/UltraData-Code)和UltraData-Math(https://huggingface.co/datasets/openbmb/UltraData-Math)。

后训练阶段,我们按三个步骤进行:SFTRLOPD。我们首先使用400B token的深度思考SFT来建立深度思考和通用聊天能力;SFT数据发布为UltraData-SFT-2605(https://huggingface.co/datasets/openbmb/UltraData-SFT-2605),智能体SFT数据发布为UltraData-SFT-Agent-2609(https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609)。然后我们为数学、代码、智能体任务、写作及相关领域训练专门的RL教师模型(相应数据也作为UltraData-RL-2609(https://huggingface.co/datasets/openbmb/UltraData-RL-2609)开源),并使用**在策略蒸馏(OPD)**将这些教师模型的能力蒸馏回一个发布模型。

MiniCPM5-2B训练方案(https://raw.githubusercontent.com/OpenBMB/MiniCPM/main/assets/minicpm5/minicpm5_2b_training_recipe.jpg)

https://huggingface.co/openbmb/MiniCPM5-2B#what-does-rl–opd-bringRL + OPD带来了什么?

RL + OPD是MiniCPM5-2B后训练的关键部分。在RL阶段,我们采用了JustRL II(https://app.notion.com/p/panhaoxuan/JustRL-II-Scaling-Small-LLMs-to-128K-Reasoning-with-a-Critic-3c77e972297c80adb8b5f4b05d267012#5f3eb56b29f048ebab5f71138f12e36f)中描述的基于评判者的算法,显著提高了训练稳定性,并在多个领域取得了显著收益。在下面列出的基准测试中,RL + OPD平均提升推理和通用能力**↑10.96分**,智能体能力**↑6.96分**。

OPD融合了通过RL训练产生的16个专家模型的能力,其中包括5个智能体专家模型。在每个响应位置,我们计算学生和教师对数概率的全词表反向KL散度作为优势估计,替代了原始的基于验证的优势。OPD直接复用用于训练每个RL教师的提示作为蒸馏数据,因此无需额外构建语料库。

MiniCPM5-2B RL + OPD收益(https://raw.githubusercontent.com/OpenBMB/MiniCPM/main/assets/minicpm5/minicpm5_2b_rl_opd_score_gains.png)

https://huggingface.co/openbmb/MiniCPM5-2B#quickstart快速开始

https://huggingface.co/openbmb/MiniCPM5-2B#vllmvLLM

pip install "vllm>=0.21" vllm serve openbmb/MiniCPM5-2B --port 8000

curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "openbmb/MiniCPM5-2B", "messages": [{"role": "user", "content": "你是谁?请简要介绍一下自己。"}], "max_tokens": 128, "temperature": 1.0 }'

https://huggingface.co/openbmb/MiniCPM5-2B#sglangSGLang

pip install "sglang[srt]>=0.5.16" python -m sglang.launch_server --model-path openbmb/MiniCPM5-2B --port 30000

curl http://localhost:30000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "openbmb/MiniCPM5-2B", "messages": [{"role": "user", "content": "你是谁?请简要介绍一下自己。"}], "max_tokens": 128, "temperature": 1.0 }'

投机解码(DSpark):我们还发布了MiniCPM5-2B-DSpark(https://huggingface.co/openbmb/MiniCPM5-2B-DSpark),这是专为MiniCPM5-2B训练的DSpark草稿模型。在SGLang中启用它可以在保持目标模型输出不变的情况下加速解码:

python -m sglang.launch_server \ --model-path openbmb/MiniCPM5-2B \ --trust-remote-code \ --speculative-algorithm DSPARK \ --speculative-draft-model-path openbmb/MiniCPM5-2B-DSpark \ --speculative-dspark-block-size 7 \ --port 30000

https://huggingface.co/openbmb/MiniCPM5-2B#transformersTransformers

pip install -U "transformers>=5.6" accelerate torch

from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "openbmb/MiniCPM5-2B" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", ) messages = [{"role": "user", "content": "你是谁?请简要介绍一下自己。"}] inputs = tokenizer.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, enable_thinking=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=128) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

推荐采样参数:temperature=1.0, top_p=0.95

https://huggingface.co/openbmb/MiniCPM5-2B#tool-calling工具调用

对于工具/函数调用,推荐使用SGLang作为后端。MiniCPM5-2B发出XML风格的工具调用,SGLang内置的minicpm5解析器会将其原生转换为OpenAI兼容的tool_calls

python -m sglang.launch_server --model-path openbmb/MiniCPM5-2B --port 30000 \ --tool-call-parser minicpm5 # 或:--tool-call-parser auto

https://huggingface.co/openbmb/MiniCPM5-2B#github-cookbooks-and-agent-skillsGitHub食谱与智能体技能

MiniCPM5-2B使用标准的LlamaForCausalLM架构,因此主流推理引擎可以直接加载它:无需自定义内核,无需模型代码分叉。有关分步部署和微调说明,请使用下面的GitHub食谱。智能体技能作为GitHub资源链接,供使用Cursor / Claude Code风格编码智能体的用户使用。

https://huggingface.co/openbmb/MiniCPM5-2B#deployment部署

https://huggingface.co/openbmb/MiniCPM5-2B#fine-tuning微调

https://huggingface.co/openbmb/MiniCPM5-2B#other-supported-frameworks其他支持的框架

除了上述部署和微调框架外,MiniCPM5-2B还受到FlagOS的多芯片部署支持。

https://huggingface.co/openbmb/MiniCPM5-2B#flagos-overviewFlagOS概述

为实现大规模部署,FlagOS提供了一种统一且可扩展的解决方案。

相似文章

MiniCPM5-1B

Reddit r/LocalLLaMA

OpenBMB 发布了 MiniCPM5-1B,这是一个密集型1B参数Transformer模型,在开源1B级模型中达到SOTA,专为设备端部署设计,支持混合推理和长上下文。

openbmb/MiniCPM-RobotManip

Hugging Face Models Trending

OpenBMB发布了MiniCPM-RobotManip,这是一个1.5B参数的视觉-语言-动作模型,用于设备端机器人操控,通过高效的流式推理和长时视觉记忆,性能超越较大模型。

MiniCPM5-1B 表明小模型竞赛尚未结束

Reddit r/ArtificialInteligence

MiniCPM5-1B 是 OpenBMB 推出的一个拥有 10 亿参数的模型,在 AIME 2025 和 τ2-Bench Telecom 上取得了令人瞩目的成绩,超越了更大的模型。它从单个检查点同时提供快速模式和推理模式,这得益于包括监督微调、强化学习和在线策略蒸馏在内的三阶段后训练过程。