双7900 xtx - 有人制作了一个针对此配置优化的llama.cpp分支,Qwen 3.8 Q8 解码速度达82 tokens/秒

Reddit r/LocalLLaMA 工具

摘要

一个针对双AMD 7900 XTX GPU优化的llama.cpp GitHub分支,显著提升了Qwen 3.8 Q8模型的解码速度至82 tokens/秒。

注意,这不是我的工作,而是我发现并想分享的东西,希望更多人能进一步推进它。https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt 我基本上在消费级PC上运行2个7900 xtx。这个仓库取qwen 3.8 Q8并优化以在此配置上运行。我的解码速度从原版llama.cpp + vulkan的约28 tokens/秒提升到约82 tokens/秒,上下文加载为60k。这很酷,我让Luna在Linux上设置整个过程,一切运行顺利。
查看原文

相似文章

I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

Reddit r/LocalLLaMA

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results: qwen 3.8 next Q3_K_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below) qwen 3.8 27B Q8_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I

双GPU llama.cpp加速

Reddit r/LocalLLaMA

llama.cpp的一个分支修复了量化KV缓存中的--split-mode tensor问题,在双GPU配置上实现高达40%的速度提升,且无质量损失。