Tag
The article describes experiments on split-K matrix multiplications, revealing that answer variability depends on block layouts and split counts, with tests on a B200 GPU showing output changes under different conditions.
Suhail reports zero availability of B200 GPUs across 13 providers, with prices heading toward $6.50-7/gpu/hr and warns that inference will get more expensive.
The article provides the optimal deployment configuration for GLM-5.2 on 8xB200 nodes, showing that NVFP4 with two TP=4 replicas achieves roughly 2x throughput over FP8 TP=8, with detailed performance data and caveats.
Claude Fable 5 achieves top results on KernelBench-Hard by hand-writing PTX code for B200 fp8 GEMM, outperforming other models and reaching 44-59% of peak performance on compute-bound shapes.