Tag
A GitHub fork of llama.cpp optimized for dual AMD 7900 XTX GPUs, significantly improving decode speed for the Qwen 3.8 Q8 model to 82 tokens per second.