@pupposandro: 2.5x faster than llama.cpp on Strix Halo. We just shipped DFlash + PFlash for the AMD Ryzen AI MAX+ 395 iGPU (gfx1151, …

X AI KOLs Following Tools

Summary

A new toolset (DFlash + PFlash) achieves 2.5x faster inference than llama.cpp on AMD Ryzen AI MAX+ 395 iGPU, demonstrating significant speedups for Qwen3.6-27B with 128 GiB unified memory.

2.5x faster than llama.cpp on Strix Halo. We just shipped DFlash + PFlash for the AMD Ryzen AI MAX+ 395 iGPU (gfx1151, 128 GiB unified memory). Qwen3.6-27B Q4_K_M, end-to-end on the same silicon: ▸ Decode: 26.85 tok/s, 2.23x faster (DFlash + DDTree, budget 22) ▸ Prefill 16K: 20.2s, 3.05x faster (PFlash) ▸ Wall clock, 16K prompt + 1K gen: 58s vs 147s ~100 GiB still free in the box. 122B and 139B MoE class is next. Massive thanks to @smpurkis0 for the contribution
Original Article

Similar Articles