Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request [P]
Summary
A monokernel approach for LLM decoding on AMD MI300X GPUs achieves up to 3,300 output tokens/s per request without speculative decoding or quantization, using memory access patterns mapped to the die topology.
Similar Articles
16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC
Describes running the GLM-5.2 model with 4-bit quantization at 12.2 tokens per second on a cluster of 16 AMD MI50 GPUs using llama.cpp's RPC, achieving coherent long-form generation at 10.7k context.
@HotAisle: This is awesome. I wonder who's MI300x they used... ;-)
Kog announces real-time LLM inference achieving 3000+ output tokens per second per request on standard datacenter GPUs, bringing high-speed inference previously limited to custom silicon to production hardware.
Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs
The article describes the optimization of AMD MI350X GPUs for running the Qwen3.6-35B-A3B LLM, achieving high output token throughput and open-sourcing the kernel to improve performance.
2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.
Packed Twin Inference (PTI) is a technique that achieves ~2× LLM throughput by running multiple token sequences in a single batch decode, exploiting weight sharing in llama.cpp without needing a draft model or additional VRAM.
Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine
Yamz Labs released Kyojin, a ROCm inference engine built on ExLlamaV3 for AMD Strix Halo, enabling two ~300B-class MoE models (GLM-5.3-Flash and MiMo-V2.6-Flash) to each fit and run on a single 128 GB mini PC, with EXL3 quantized weights hitting up to 580 tok/s prefill and 44 tok/s speculative decode.