First Kimi K3 results on home lab ~ 4t/s

Reddit r/LocalLLaMA Models

Summary

First benchmark results for Kimi K3 on a home lab setup show inference speed of approximately 4 tokens per second.

No content available
Original Article

Similar Articles

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

Hacker News Top

WASTE is a new open-source C inference engine that streams expert weights from disk to run the 2.78-trillion-parameter Kimi K3 model on a consumer laptop with just 29 GB of RAM, achieving 0.49–0.54 tokens/s.

Kimi K3 Benchmarks

Reddit r/singularity

Kimi K3 has achieved notable results in recent AI benchmarks, showcasing its capabilities.

I got Kimi-k3 running.....

Reddit r/LocalLLaMA

User successfully runs the Kimi-k3 model using llama.cpp on high-end hardware, achieving low tokens per second (0.41 prompt eval, 0.23 generation).

Kimi K3 Coding Benchmarks

Reddit r/singularity

Kimi K3 coding benchmarks article discussing performance of the Kimi K3 model on coding tasks.