@ciprianveg: Great speed-up: Full Kimi K3 — the 2.8T-parameter model running on 16× NVIDIA GB10 (Spark) - V5 coming by tomorrow, cca…
Summary
Kimi K3, a 2.8T-parameter AI model, is running on a cluster of 16 NVIDIA GB10s with impressive speed metrics, and a V5 update is expected tomorrow with about 20% faster performance and improved concurrency.
View Cached Full Text
Cached at: 09/20/26, 05:16 AM
Great speed-up: Full Kimi K3 — the 2.8T-parameter model running on 16× NVIDIA GB10 (Spark) - V5 coming by tomorrow, cca 20% faster and much improved concurrency: (~30t/s C1, 87t/s C8 ) 136t/s peak speed at C8:
Coding: game-bench 1 2 4 8 (3000-token runs) C1: 29.81 tok/s C2: 42.00 agg / 21.00 mean stream C4: 58.00 agg / 14.50 mean stream C8: 87.12 agg / 10.89 mean stream Prose: llama-bench coherent corpus modeltestt/speak t/s KIMI-K3pp2048 @ d4000880.35 ± 0.00 KIMI-K3tg2048 @ d400023.59 ± 0.0046.00 ± 0.00 KIMI-K3pp2048 @ d100000812.56 ± 0.00 KIMI-K3tg2048 @ d10000021.13 ± 0.0033.00 ± 0.00 KIMI-K3pp2048 @ d200000848.73 ± 0.00 KIMI-K3tg2048 @ d20000020.03 ± 0.0038.00 ± 0.00 https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174/141?u=ciprianveg…
Similar Articles
Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model runs on a 16x GB10 cluster at 20+ tokens per second average, with plans to publish the vllm image and instructions.
@UnTalNixon_exe: Forget about GPUs and million-dollar clusters. They just made the world's largest open model (Kimi K3 – 2.78 trillion p…
A new tool called kimi-k3-in-c runs the 2.78T-parameter Kimi K3 open model on a single CPU with as little as 8.24 GB RAM, streaming experts from disk and achieving deterministic output at 10-32 seconds per token.
@levie: The k3 weights have arrived
Kimi.ai released the model weights and technical report for Kimi K3, a 2.8T parameter MoE model with native visual understanding and a 1M-token context window, claiming 2.5x intelligence per unit of compute.
@svpino: I've been testing Kimi K3, and holy smokes, this is the best open-weight model the world has seen. It's a 2.8T-paramete…
Santiago Valdarrama reports that Kimi K3 is the best open-weight model he has tested, a 2.8T-parameter vision model with tool calling, reasoning, and a 1M context window.
@thealexker: underrated gems in Kimi-K3 release: > an early K3 wrote the majority of the kernels in the late development stages > it…
Kimi.ai released Kimi K3, a 2.8 trillion parameter multimodal model with 1 million context, featuring novel Delta Attention and Attention Residuals, and a self-optimizing stack including MiniTriton compiler. The model achieves up to 6.3x faster decoding and ~25% higher training efficiency.