llama.cpp 自适应 MTP PR#27210

Reddit r/LocalLLaMA 工具

摘要

一个用于自适应 MTP 的拉取请求 (PR#27210) 已提交到 llama.cpp,该实现是用于LLM推理的C/C++代码,设置简单且性能高效。

暂无内容
查看原文
查看缓存全文

缓存时间: 2026/08/17 20:23

ggml-org/llama.cpp

源码:https://github.com/ggml-org/llama.cpp

llama.cpp

使用 C/C++ 实现的 llama LLM 推理引擎
许可证:MIT (https://opensource.org/licenses/MIT)
发行版 (https://github.com/ggml-org/llama.cpp/releases)
服务器 (https://github.com/ggml-org/llama.cpp/actions/workflows/server.yml)
Docker (https://github.com/ggml-org/llama.cpp/actions/workflows/docker.yml)
Winget (https://github.com/ggml-org/llama.cpp/actions/workflows/winget.yml)
宣言 (https://github.com/ggml-org/llama.cpp/discussions/205)
/ ggml (https://github.com/ggml-org/ggml)
/ 操作 (https://github.com/ggml-org/llama.cpp/blob/master/docs/ops.md)
/ 维护者 PR (https://github.com/ggml-org/llama.cpp/issues?q=is%3Apr%20is%3Aopen%20draft%3AFalse%20(author%3Argerganov%20OR%20author%3AKitaitiMakoto%20OR%20author%3Adanbev%20OR%20author%3Aaldehir%20OR%20author%3Amax-krasnyansky%20OR%20author%3ACISC%20OR%20author%3Aggerganov%20OR%20author%3Aam17an%20OR%20author%3Abartowski1182%20OR%20author%3Ahipudding%20OR%20author%3AServeurpersoCom%20OR%20author%3Apwilkin%20OR%20author%3Areeselevine%20OR%20author%3Angxson%20OR%20author%3Ajeffbolznv%20OR%20author%3A0cc4m%20OR%20author%3Aangt%20OR%20author%3AIMbackK%20OR%20author%3Aarthw%20OR%20author%3AJohannesGaessler%20OR%20author%3AORippler%20OR%20author%3Aruixiang63%20OR%20author%3Axctan%20OR%20author%3Aallozaur%20OR%20author%3Ayomaytk%20OR%20author%3Aaendk%20OR%20author%3Agaugarg-nv%20OR%20author%3Ataronaeo%20OR%20author%3Aforforever73%20OR%20author%3Alhez%20OR%20author%3Anetrunnereve%20OR%20author%3Afairydreaming)%20sort%3Aupdated-desc)
/ 编译时间 (https://github.com/ggml-org/llama.cpp-dev/blob/master/README-compile-times.md)
/ llama API 库 (https://github.com/ggml-org/llama.cpp/issues/9289)
/ llama-server REST API (https://github.com/ggml-org/llama.cpp/issues/9291)

快速开始

在您的机器上安装 llama.cpp 有几种方式:

  • 访问 https://llama.app 并遵循说明
  • 使用 Docker 运行 - 参阅我们的 Docker 文档
  • 从发行版页面 (https://github.com/ggml-org/llama.cpp/releases) 下载预构建二进制文件
  • 克隆此仓库从源代码构建 - 查看 我们的构建指南

安装完成后:

# 从 Hugging Face 直接下载并运行模型  
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF  

# 启动兼容 OpenAI 的 API 服务器  
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF  

使用 llama cli 进行 VLM 会话
基于 llama serve 的内置 Web UI

描述

llama.cpp 的主要目标是以最少的设置,在本地和云端的广泛硬件上实现具有顶尖性能的 LLM(和 VLM)推理。

  • 不依赖任何库的纯 C/C++ 实现
  • Apple 硅芯片是一等公民 - 通过 ARM NEON、Accelerate 和 Metal 框架优化
  • 支持 x86 架构的 AVX、AVX2、AVX512 和 AMX
  • 支持 RISC-V 架构的 RVV、ZVFH、ZFH、ZICBOP 和 ZIHINTPAUSE
  • 支持 1.5 位、2 位、3 位、4 位、5 位、6 位和 8 位整数量化,以加快推理速度并减少内存使用
  • 为在 NVIDIA GPU 上运行 LLM 提供自定义 CUDA 内核(通过 HIP 支持 AMD GPU,通过 MUSA 支持 Moore Threads GPU)
  • Vulkan 和 SYCL 后端支持
  • CPU+GPU 混合推理,可部分加速超过总显存容量的模型

llama.cpp 项目基于 ggml (https://github.com/ggml-org/ggml) 库构建。

支持的后端

后端目标设备
BLAS所有设备
BLIS所有设备
CANNAscend NPU
CUDANvidia GPU
HIPAMD GPU
Hexagon [进行中]Snapdragon
IBM zDNNIBM Z & LinuxONE
MUSAMoore Threads GPU
MetalApple 硅芯片
OpenCLAdreno GPU
OpenVINO [进行中]Intel CPU、GPU 和 NPU
RPC (https://github.com/ggml-org/llama.cpp/tree/master/tools/rpc)所有设备
SYCLIntel GPU
VirtGPUVirtGPU APIR
VulkanGPU
WebGPU所有设备
ZenDNNAMD CPU

文档

工具

开发

贡献

  • 贡献者可以提交 PR
  • 根据贡献情况邀请协作者
  • 维护者可以将代码推送到 llama.cpp 仓库的分支,并将 PR 合并到 master 分支
  • 非常感谢对管理 issues、PR 和项目的任何帮助!
  • 阅读 CONTRIBUTING.md 了解更多信息

致谢

  • yhirose/cpp-httplib (https://github.com/yhirose/cpp-httplib)
    • 单头文件 HTTP 服务器,被 llama-server 使用
    • MIT 许可证
  • stb-image (https://github.com/nothings/stb)
    • 单头文件图像格式解码器,被多模态子系统使用
    • 公共领域
  • nlohmann/json (https://github.com/nlohmann/json)
    • 单头文件 JSON 库,被各种工具/示例使用
    • MIT 许可证
  • miniaudio.h (https://github.com/mackron/miniaudio)
    • 单头文件音频格式解码器,被多模态子系统使用
    • 公共领域
  • subprocess.h (https://github.com/sheredom/subprocess.h)
    • 用于 C 和 C++ 的单头文件进程启动解决方案
    • 公共领域

相似文章

MTP PR 已合并!!!

Reddit r/LocalLLaMA

与 LLaMA 模型相关的 MTP(可能指模型训练管道或类似内容)拉取请求已合并,标志着一个里程碑。