@jun_song: Working on fitting Kimi-K2.6 (1T) on 128GB Mac. Trying to get 40tok/s, and minimize the quality loss.

X AI KOLs Timeline News

Summary

A developer is optimizing the Kimi-K2.6 (1T) model to run efficiently on a 128GB Mac, targeting 40 tokens per second while minimizing quality loss.

Working on fitting Kimi-K2.6 (1T) on 128GB Mac. Trying to get 40tok/s, and minimize the quality loss.
Original Article
View Cached Full Text

Cached at: 05/11/26, 12:42 PM

Working on fitting Kimi-K2.6 (1T) on 128GB Mac.

Trying to get 40tok/s, and minimize the quality loss.

Similar Articles

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

Hacker News Top

WASTE is a new open-source C inference engine that streams expert weights from disk to run the 2.78-trillion-parameter Kimi K3 model on a consumer laptop with just 29 GB of RAM, achieving 0.49–0.54 tokens/s.