dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode

Reddit r/LocalLLaMA Tools

Summary

A GitHub fork of llama.cpp optimized for dual AMD 7900 XTX GPUs, significantly improving decode speed for the Qwen 3.8 Q8 model to 82 tokens per second.

Note not my work, but something i found and wanted to share so hopefully more people can push this along even further. https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt I basically run 2 x 7900 xtx on a consumer pc. This repo takes qwen 3.8 Q8 and optimizes it to run on this setup. My decode went from about 28 tokens / seconds on vanila lamacpp + vulkan to about 82 tokens / second at 60k context load. its pretty neat i had luna setup the whole thing in linux and it worked like a charm.
Original Article

Similar Articles

I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

Reddit r/LocalLLaMA

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results: qwen 3.8 next Q3_K_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below) qwen 3.8 27B Q8_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I

Dual GPU llama.cpp speedup

Reddit r/LocalLLaMA

A fork of llama.cpp fixes the --split-mode tensor issue with quantized KV caches, achieving up to 40% speed improvement on dual GPU setups without quality loss.