@QuixiAI: @Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test

X AI KOLs Following Models

Summary

Kimi K2.6 achieves 56 tokens per second on a single MI300X GPU; user plans further throughput benchmarking.

@Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Original Article
View Cached Full Text

Cached at: 04/21/26, 01:02 PM

@Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test

Similar Articles

First Kimi K3 results on home lab ~ 4t/s

Reddit r/LocalLLaMA

User shares first home-lab results running Kimi K3 on 768GB DDR5 and 2x5090 using a llama.cpp fork and Q2_K quant, reporting prefill speeds of 50-70 tps and decoding tps that increases over time.