@jun_song: Working on fitting Kimi-K2.6 (1T) on 128GB Mac. Trying to get 40tok/s, and minimize the quality loss.
Summary
A developer is optimizing the Kimi-K2.6 (1T) model to run efficiently on a 128GB Mac, targeting 40 tokens per second while minimizing quality loss.
View Cached Full Text
Cached at: 05/11/26, 12:42 PM
Working on fitting Kimi-K2.6 (1T) on 128GB Mac.
Trying to get 40tok/s, and minimize the quality loss.
Similar Articles
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
WASTE is a new open-source C inference engine that streams expert weights from disk to run the 2.78-trillion-parameter Kimi K3 model on a consumer laptop with just 29 GB of RAM, achieving 0.49–0.54 tokens/s.
@QuixiAI: @Kimi_Moonshot K2.6 running on my mi300x, 56 tps (single request). I will run a throughput test
Kimi K2.6 achieves 56 tokens per second on a single MI300X GPU; user plans further throughput benchmarking.
@HotAisle: Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → …
Kimi K2.6 paired with DFlash inference system achieves 508 tokens/s on 8×AMD MI300X, a 5.6× throughput jump from 90 tokens/s baseline with zero quality loss.
Kimi K3 full model running on 16x GB10 cluster at 20+tps
Kimi K3 full model runs on a 16x GB10 cluster at 20+ tokens per second average, with plans to publish the vllm image and instructions.
@UnTalNixon_exe: Forget about GPUs and million-dollar clusters. They just made the world's largest open model (Kimi K3 – 2.78 trillion p…
A new tool called kimi-k3-in-c runs the 2.78T-parameter Kimi K3 open model on a single CPU with as little as 8.24 GB RAM, streaming experts from disk and achieving deterministic output at 10-32 seconds per token.