Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

Reddit r/LocalLLaMA News

Summary

User reports running Qwen 3.6 35B A3B-Q8_0 gguf on a Radeon 7600 with llama.cpp and ROCm, achieving 21 tokens per second after VRAM overclocking, with a note about a display-related performance bug.

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 \ --n-cpu-moe 36 \ --no-mmap \ -ctk q8_0 \ -ctv q8_0 \ -fa 1 \ -c 9000 \ So rebuilt my llama.cpp build to run rocm 7.14 tokens increased to upper 19 token/per second then overclocked the vram to the maximum LACTL will allow now 21 token/s also. Weird bug if I am watching the tokens being generated by llama it drops to 13 tokens per second but window minimized it goes up to 21 tokens per second weird.
Original Article

Similar Articles