low-bpw

Tag

Cards List
#low-bpw

Anyone tested the IQ1_M 342GB Pruned Kimi K3? Is it usable?

Reddit r/LocalLLaMA · 3d ago Cached

This is a highly experimental GGUF version of the 2.8T-parameter Kimi K3 MoE model, with 55% of experts pruned and quantized to ~2.15 bpw (319 GiB). It requires a specific llama.cpp PR and custom patches to run, and includes detailed instructions for usage.

0 favorites 0 likes
← Back to home

Submit Feedback