New Unsloth KImi K3 drops! Q1_0 (466GB), TQ1_0(509GB), IQ1_M(649),TQ2_0(551GB)!!

Reddit r/LocalLLaMA Models

Summary

Unsloth releases new GGUF quantizations of Kimi K3, ranging from 466GB to 649GB, enabling efficient deployment of the large model.

The smallest UD-Q1_0 is 466GB, TQ2_0 551GB. Well done team Unsloth! https://huggingface.co/unsloth/Kimi-K3-GGUF
Original Article

Similar Articles

Anyone tried the Q1 Kimi K3 yet? (555GB)

Reddit r/LocalLLaMA

Kimi K3 is a massive 2.9 trillion parameter mixture-of-experts model with 104B active parameters, 1M context length, and native MXFP4 training, now available in GGUF quantizations ranging from 540GB to smaller sizes, though requiring substantial hardware to run.

Kimi K2.6 Unsloth GGUF is out

Reddit r/LocalLLaMA

Unsloth has released a GGUF-quantized version of the Kimi K2.6 model, enabling efficient local inference.

Anyone tested the IQ1_M 342GB Pruned Kimi K3? Is it usable?

Reddit r/LocalLLaMA

This is a highly experimental GGUF version of the 2.8T-parameter Kimi K3 MoE model, with 55% of experts pruned and quantized to ~2.15 bpw (319 GiB). It requires a specific llama.cpp PR and custom patches to run, and includes detailed instructions for usage.

unsloth/Kimi-K2.6-GGUF

Hugging Face Models Trending

Unsloth releases quantized GGUF versions of the open-source 1T-parameter Kimi K2.6 MoE model, optimized for long-horizon coding, autonomous agent swarms, and production-ready design tasks.