First Kimi K3 results on home lab ~ 4t/s
Summary
First benchmark results for Kimi K3 on a home lab setup show inference speed of approximately 4 tokens per second.
Similar Articles
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
WASTE is a new open-source C inference engine that streams expert weights from disk to run the 2.78-trillion-parameter Kimi K3 model on a consumer laptop with just 29 GB of RAM, achieving 0.49–0.54 tokens/s.
Kimi K3 Benchmarks
Kimi K3 has achieved notable results in recent AI benchmarks, showcasing its capabilities.
I got Kimi-k3 running.....
User successfully runs the Kimi-k3 model using llama.cpp on high-end hardware, achieving low tokens per second (0.41 prompt eval, 0.23 generation).
@HotAisle: Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → …
Kimi K2.6 paired with DFlash inference system achieves 508 tokens/s on 8×AMD MI300X, a 5.6× throughput jump from 90 tokens/s baseline with zero quality loss.
Kimi K3 Coding Benchmarks
Kimi K3 coding benchmarks article discussing performance of the Kimi K3 model on coding tasks.