@wafer_ai: BREAKING: these engineers figured out how to serve GLM 5.2 on @AMD MI355X at 2626 tok/s/node and 213 tok/s single strea…
Summary
Engineers successfully serve GLM 5.2 on AMD MI355X at 2626 tok/s per node and 213 tok/s single stream, achieving ~80% of B200 throughput at over 2x lower cost than Blackwell.
View Cached Full Text
Cached at: 07/04/26, 08:41 AM
🚨 BREAKING:
these engineers figured out how to serve GLM 5.2 on @AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell
that’s ~80% of B200 throughput at over 2x lower cost
full write-up in reply to see how https://t.co/1wAickkQEm
BREAKING:
these engineers figured out how to serve GLM 5.2 on @AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell
that’s ~80% of B200 throughput at over 2x lower cost
full write-up in reply to see how
this is how they did it:
yo that’s my dad you’re talking about
lollll
not as cool as you @tomgreenwald
thank you!
thank you!
we’re always on time
Similar Articles
Performance per dollar is getting faster and cheaper
Wafer demonstrates that AMD MI355X GPUs offer competitive inference performance for frontier models like GLM5.2 at significantly lower cost than NVIDIA Blackwell, achieving 80% of B200 throughput at under half the price, using MXFP4 quantization and sglang.
@wafer_ai: BREAKING We're on Hacker News again we figured out how to serve Kimi K3 at 3.8x higher throughput and 71% lower cost on…
Wafer announces that it can serve Kimi K3 on AMD MI355X at 3.8x higher throughput and 71% lower cost than on B200 nodes, arguing that AMD's large VRAM and software support make it the best performance-per-dollar choice for frontier models.
16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC
Describes running the GLM-5.2 model with 4-bit quantization at 12.2 tokens per second on a cluster of 16 AMD MI50 GPUs using llama.cpp's RPC, achieving coherent long-form generation at 10.7k context.
I did some model hacks, and got GLM5.2 from about 2.5 tok/s to >50 tok/s on my GH200 system.
A detailed blog post describing how to dramatically speed up GLM-5.2 inference on a dual Grace Hopper system from 2.5 tok/s to over 50 tok/s by stopping model cross-module traffic and grafting an FP8 MTP head onto the INT4 base.
LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels
LFM2.5 230M model achieves 1,400 tokens per second in-browser using custom WebGPU kernels, demonstrating efficient local inference.