A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.
Summary
A user runs Kimi K3, a 2.8T-parameter open-weight MoE model, on 80 RTX 5090 GPUs using only GDDR7 and Ethernet—no HBM—achieving 20 tok/s single stream, the first frontier model to run on consumer hardware.
View Cached Full Text
Cached at: 07/28/26, 06:35 AM
we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s.
20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb.
A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized.
The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it.
@Kimi_Moonshot
Kimi.ai (@Kimi_Moonshot): Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside
Similar Articles
Kimi K3 is like an F1 machine inside a show window.
Moonshot has released Kimi K3, an extremely large AI model that is virtually impossible to run on local workstations even with multiple RTX 6000 Blackwell GPUs, prompting the author to seek ways to run it via hacking or distillation.
@TheAhmadOsman: RTX 3090 owners tonight will be running Kimi_K3_3T_Q_0.001_K GGUF
A quantized GGUF version of the Kimi K3 model is now available, optimized for running on RTX 3090 GPUs.
@QuixiAI: @Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Kimi K2.6 achieves 56 tokens per second on a single MI300X GPU; user plans further throughput benchmarking.
I got Kimi-k3 running.....
User successfully runs the Kimi-k3 model using llama.cpp on high-end hardware, achieving low tokens per second (0.41 prompt eval, 0.23 generation).
VRAM disk cache of MoE makes 340 pp/s 9.6 tg/s for Kimi 2.7 on a single dgx spark
A detailed strategy using VRAM as disk cache with unified memory and llama.cpp settings achieves 340 pp/s and 9.6 tg/s for Kimi K2.7 on a single DGX Spark.