A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.

Reddit r/LocalLLaMA News

Summary

A user runs Kimi K3, a 2.8T-parameter open-weight MoE model, on 80 RTX 5090 GPUs using only GDDR7 and Ethernet—no HBM—achieving 20 tok/s single stream, the first frontier model to run on consumer hardware.

we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s. 20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb. A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized. The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it. @Kimi_Moonshot
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:35 AM

we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s.

20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb.

A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized.

The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it.

@Kimi_Moonshot

Kimi.ai (@Kimi_Moonshot): Releasing the model weights and technical report of Kimi K3.

Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.

New model architecture: 2.5x the intelligence per unit of compute, not just more params.

Alongside

Similar Articles

Kimi K3 is like an F1 machine inside a show window.

Reddit r/LocalLLaMA

Moonshot has released Kimi K3, an extremely large AI model that is virtually impossible to run on local workstations even with multiple RTX 6000 Blackwell GPUs, prompting the author to seek ways to run it via hacking or distillation.

I got Kimi-k3 running.....

Reddit r/LocalLLaMA

User successfully runs the Kimi-k3 model using llama.cpp on high-end hardware, achieving low tokens per second (0.41 prompt eval, 0.23 generation).