Release b11003 · ggml-org/llama.cpp

Reddit r/LocalLLaMA 工具

摘要

llama.cpp 发布版本 b11003,这是一个 C/C++ 工具,可在各种硬件上以最少的设置实现高效的大语言模型推理。

暂无内容
查看原文
查看缓存全文

缓存时间: 2026/09/16 17:35

ggml-org/llama.cpp 来源:https://github.com/ggml-org/llama.cpp # llama.cpp C/C++ 语言实现的 llama 大语言模型推理 许可证:MIT (https://opensource.org/licenses/MIT) 正式发布版 (https://github.com/ggml-org/llama.cpp/releases?q=tag:v0) 每日构建版 (https://github.com/ggml-org/llama.cpp/releases?q=b) 服务器 (https://github.com/ggml-org/llama.cpp/actions/workflows/server.yml) Docker (https://github.com/ggml-org/llama.cpp/actions/workflows/docker.yml) Winget (https://github.com/ggml-org/llama.cpp/actions/workflows/winget.yml) ggml (https://github.com/ggml-org/ggml) / 运算操作 (https://github.com/ggml-org/llama.cpp/blob/master/docs/ops.md) / 维护者PR (https://github.com/ggml-org/llama.cpp/issues?q=is%3Apr%20is%3Aopen%20draft%3AFalse%20(author%3Argerganov%20OR%20author%3AKitaitiMakoto%20OR%20author%3Adanbev%20OR%20author%3Aaldehir%20OR%20author%3Amax-krasnyansky%20OR%20author%3ACISC%20OR%20author%3Aggerganov%20OR%20author%3Aam17an%20OR%20author%3Ajhen0409%20OR%20author%3Abartowski1182%20OR%20author%3Anikwen%20OR%20author%3Ahipudding%20OR%20author%3Aravi9%20OR%20author%3AServeurpersoCom%20OR%20author%3Apwilkin%20OR%20author%3Areeselevine%20OR%20author%3Angxson%20OR%20author%3Ajeffbolznv%20OR%20author%3Amarty1885%20OR%20author%3A0cc4m%20OR%20author%3ATitaniumtown%20OR%20author%3Aangt%20OR%20author%3AIMbackK%20OR%20author%3Aarthw%20OR%20author%3AJohannesGaessler%20OR%20author%3AORippler%20OR%20author%3Aruixiang63%20OR%20author%3Axctan%20OR%20author%3Aallozaur%20OR%20author%3Ayomaytk%20OR%20author%3Aaendk%20OR%20author%3Awine99%20OR%20author%3Agaugarg-nv%20OR%20author%3Ataronaeo%20OR%20author%3Aforforever73%20OR%20author%3Alhez%20OR%20author%3Anetrunnereve%20OR%20author%3Afairydreaming)%20sort%3Aupdated-desc) / 开发统计 (https://github.com/ggml-org/llama.cpp-dev) / lib llama API (https://github.com/ggml-org/llama.cpp/issues/9289) / llama-server REST API (https://github.com/ggml-org/llama.cpp/issues/9291) ## 快速上手 几种将 llama.cpp 安装到您机器上的方法: - 访问 https://llama.app 并按照说明操作 - 使用 Docker 运行 - 请参阅我们的 Docker 文档 - 从发布页面下载预编译二进制文件 (https://github.com/ggml-org/llama.cpp/releases) - 通过克隆本仓库从源代码构建 - 请查看我们的构建指南 安装完成后: sh # 直接从 Hugging Face 下载并运行模型 llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF # 启动兼容 OpenAI 接口的 API 服务器 llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF 使用 llama cli 的 VLM 会话 针对 llama serve 的内置 Web UI ## 描述 llama.cpp 的主要目标是以最少的设置和顶尖的性能,在多种硬件上(本地和云端)实现大语言模型(及视觉语言模型)推理。 - 纯 C/C++ 实现,无任何依赖 - Apple 芯片得到一等支持 - 通过 ARM NEON、Accelerate 和 Metal 框架进行优化 - 支持 x86 架构的 AVX、AVX2、AVX512 和 AMX 指令集 - 支持 RISC-V 架构的 RVV、ZVFH、ZFH、ZICBOP 和 ZIHINTPAUSE 扩展 - 支持 1.5 位、2 位、3 位、4 位、5 位、6 位和 8 位整数量化,以加快推理速度并减少内存使用 - 为在 NVIDIA GPU 上运行大语言模型提供定制 CUDA 内核(通过 HIP 支持 AMD GPU,通过 MUSA 支持 Moore Threads GPU) - 支持 Vulkan 和 SYCL 后端 - 支持 CPU+GPU 混合推理,可部分加速显存容量不足的大型模型 llama.cpp 项目构建于 ggml (https://github.com/ggml-org/ggml) 库之上。 ## 支持的后端 | 后端 | 目标设备 | | — | — | | BLAS | 所有设备 | | BLIS | 所有设备 | | CANN | 昇腾 NPU | | CUDA | Nvidia GPU | | HIP | AMD GPU | | Hexagon | Snapdragon | | IBM zDNN | IBM Z & LinuxONE | | MUSA | Moore Threads GPU | | Metal | Apple 芯片 | | OpenCL | Adreno GPU | | OpenVINO [进行中] | Intel CPU、GPU 和 NPU | | RPC (https://github.com/ggml-org/llama.cpp/tree/master/tools/rpc) | 所有设备 | | SYCL | Intel GPU | | VirtGPU | VirtGPU APIR | | Vulkan | GPU | | WebGPU | 所有设备 | | ZenDNN | AMD CPU | ## 文档 #### 工具 - cli - 补全 - 服务器 - GBNF 语法 #### 开发 - 如何构建 - 在 Docker 上运行 - 在 Android 上构建 - 多 GPU 使用 - 性能问题排查 - GGML 小贴士与技巧 (https://github.com/ggml-org/llama.cpp/wiki/GGML-Tips-&-Tricks) - XCFramework - 补全功能 - 模型 - 发布流程 ## 贡献指南 - 贡献者可以提交 PR - 根据贡献情况,协作者将被邀请加入 - 维护者可以向 llama.cpp 仓库的分支推送代码,并将 PR 合并到 master 分支 - 非常感谢任何在管理问题、PR 和项目方面的帮助! - 请阅读 CONTRIBUTING.md 了解更多信息 ## 致谢 - yhirose/cpp-httplib (https://github.com/yhirose/cpp-httplib) - 单头文件 HTTP 服务器,被 llama-server 使用 - MIT 许可证 - nothings/stb (https://github.com/nothings/stb) - 单头文件图像格式解码器,被多模态子系统使用 - 公共领域 - nlohmann/json (https://github.com/nlohmann/json) - 单头文件 JSON 库,被各种工具/示例使用 - MIT 许可证 - mackron/miniaudio (https://github.com/mackron/miniaudio) - 单头文件音频格式解码器,被多模态子系统使用 - 公共领域 - sheredom/subprocess.h (https://github.com/sheredom/subprocess.h) - C 和 C++ 的单头文件进程启动解决方案 - 公共领域

相似文章

ggml-org/llama.cpp

GitHub Trending (daily)

llama.cpp 是一个开源 C/C++ 库,用于在本地硬件上高效运行 LLM 推理,支持多种量化方法和多后端(CPU、GPU 等)。

Llama.cpp v0.1.0

Hacker News Top

Llama.cpp v0.1.0 是用于高效 LLM 和 VLM 推理的 C/C++ 实现,支持广泛的硬件,配置简单且性能卓越。

llama.cpp 里程碑

Reddit r/LocalLLaMA

llama.cpp 发布里程碑版本,该开源库可在消费级硬件上本地运行大型语言模型,本次更新提升了性能或增加了新功能。