I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
Summary
The article describes hosting the Kimi K3 AI model with 2.8 trillion parameters using 8 B300 GPUs, achieving 92 tokens per second and costing $190 per million tokens, while comparing it with Unsloth's dynamic GGUF quantization method.
Similar Articles
Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model runs on a 16x GB10 cluster at 20+ tokens per second average, with plans to publish the vllm image and instructions.
@UnTalNixon_exe: Forget about GPUs and million-dollar clusters. They just made the world's largest open model (Kimi K3 – 2.78 trillion p…
A new tool called kimi-k3-in-c runs the 2.78T-parameter Kimi K3 open model on a single CPU with as little as 8.24 GB RAM, streaming experts from disk and achieving deterministic output at 10-32 seconds per token.
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
WASTE is a new open-source C inference engine that streams expert weights from disk to run the 2.78-trillion-parameter Kimi K3 model on a consumer laptop with just 29 GB of RAM, achieving 0.49–0.54 tokens/s.
@QuixiAI: @Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Kimi K2.6 achieves 56 tokens per second on a single MI300X GPU; user plans further throughput benchmarking.
Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
An updated benchmark shows self-hosting Kimi K3 on 8×B300 nodes achieves 86.4% task resolution at roughly 20% higher hardware cost compared to GLM-5.2 on B200 nodes, though with lower throughput.