@QuixiAI: @Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Summary
Kimi K2.6 achieves 56 tokens per second on a single MI300X GPU; user plans further throughput benchmarking.
View Cached Full Text
Cached at: 04/21/26, 01:02 PM
@Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Similar Articles
@HotAisle: Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → …
Kimi K2.6 paired with DFlash inference system achieves 508 tokens/s on 8×AMD MI300X, a 5.6× throughput jump from 90 tokens/s baseline with zero quality loss.
I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
The article describes hosting the Kimi K3 AI model with 2.8 trillion parameters using 8 B300 GPUs, achieving 92 tokens per second and costing $190 per million tokens, while comparing it with Unsloth's dynamic GGUF quantization method.
Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model runs on a 16x GB10 cluster at 20+ tokens per second average, with plans to publish the vllm image and instructions.
@jun_song: Working on fitting Kimi-K2.6 (1T) on 128GB Mac. Trying to get 40tok/s, and minimize the quality loss.
A developer is optimizing the Kimi-K2.6 (1T) model to run efficiently on a 128GB Mac, targeting 40 tokens per second while minimizing quality loss.
First Kimi K3 results on home lab ~ 4t/s
User shares first home-lab results running Kimi K3 on 768GB DDR5 and 2x5090 using a llama.cpp fork and Q2_K quant, reporting prefill speeds of 50-70 tps and decoding tps that increases over time.