双7900 xtx - 有人制作了一个针对此配置优化的llama.cpp分支,Qwen 3.8 Q8 解码速度达82 tokens/秒
摘要
一个针对双AMD 7900 XTX GPU优化的llama.cpp GitHub分支,显著提升了Qwen 3.8 Q8模型的解码速度至82 tokens/秒。
相似文章
I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)
I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results: qwen 3.8 next Q3_K_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below) qwen 3.8 27B Q8_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I
使用我的新INT8 vLLM分支,在4x MI100配置(成本$6.5k)上实现Qwen3.8 27B C8的972 TG / 5,680 PP性能
一个新的vLLM分支为较旧的AMD MI100 GPU上的Qwen3.8 27B引入全面的INT8优化,实现高达972 tokens/秒的吞吐量,并进行严格的精度验证。
双Radeon R9700——在llama.cpp上运行Qwen 3.6 27B Q8 MTP
关于在使用ROCm的llama.cpp上,于双AMD Radeon R9700配置下运行Qwen 3.6 27B Q8模型的技术报告,包括性能基准测试和配置详情。
Qwen3.8-Flash-Next 在双 RTX 3090 + DDR4 上:解码速度从 17 提升至 25-29 令牌/秒,使用专家缓存 PR
用户基准测试并详述了 llama.cpp 中的 GPU 常驻专家缓存 PR,该 PR 将 Qwen3.8-Flash-Next 在双 RTX 3090 系统上的解码速度从 17 提升至 25-29 令牌/秒。
双GPU llama.cpp加速
llama.cpp的一个分支修复了量化KV缓存中的--split-mode tensor问题,在双GPU配置上实现高达40%的速度提升,且无质量损失。