If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

Reddit r/LocalLLaMA Tools

Summary

A user recommends a GitHub project for faster inference of Qwen 3.8 Flash Next on Strix Halo, showing benchmark speeds that significantly outperform llama.cpp.

Here's the project. I have nothing to do with it. I'm just an amazed user. https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BENCHMARKS.md Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat. "[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]" That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp. It works with models other than QFN, but the current number is small. You can find the list on their project page.
Original Article

Similar Articles

Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context

Reddit r/LocalLLaMA

The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.