@bnjmn_marie: For LFM2.5 8B A1B, the MoQ GGUFs are the best They have the best ratio accuracy/size Again, it's interesting to see the…
Summary
The poster states that the MoQ GGUFs of the LFM2.5 8B A1B model offer the best accuracy-to-size ratio, advising against using versions with less than 95% accuracy recovery.
View Cached Full Text
Cached at: 06/10/26, 09:47 AM
For LFM2.5 8B A1B, the MoQ GGUFs are the best
They have the best ratio accuracy/size
Again, it’s interesting to see the performance difference for low-bit versions, but I wouldn’t use any version that doesn’t achieve a minimum of 95% accuracy recovery. https://t.co/GfXB2re4xQ
Similar Articles
@WaleedAhmad1a10: Check out the Qwen 3.5 27B MoQ GGUFs :
A Hugging Face repository (kaitchup/Qwen3.6-27B-GGUF-MoQ) provides GGUF quantized weights for the Qwen3.6-27B MoQ model, enabling local inference with tools like llama.cpp and Ollama.
@aisearchio: GLM 5.2 GGUF is already here! 8-bit is ~half the size of the full model. Smaller versions coming soon https://huggingfa…
GLM 5.2 GGUF quantized model is released, with 8-bit version half the size of the full model; smaller versions are coming soon.
Calibrating 2-bit GGUFs (<10Gb) for agentic coding tasks
This article introduces calibrated 2-bit GGUF quantizations of the Qwopus3.6-27B-Coder model for agentic coding tasks, demonstrating that the IQ2_M quant (9.74 GiB) achieves a 63% pass rate on the SWE-rebench benchmark, comparable to a Q5_K_M quant at half the size.
@TeksEdge: Unsloth released the fastest Qwen3.6-27B MTP GGUF I've tested. Time to upgrade. Compared to the previous GGUF, Q4/Q6 XL…
Unsloth has released an optimized GGUF version of the Qwen3.6-27B MTP model, achieving significantly faster inference speeds (up to 114 tok/s on an RTX 5090) compared to previous quantizations.
Gemma 4 26B-A4B GGUF Benchmarks
Unsloth has released KL Divergence benchmarks for Gemma 4 26B-A4B GGUF quantizations, showing Unsloth GGUFs top 21 of 22 sizes on the Pareto frontier. They also introduced a new UD-IQ4_NL_XL quant fitting in 16GB VRAM and updated Q6_K and MLX quants for both Gemma 4 and Qwen3.6.