Qwen 27B 3.8 quants: How low can you go?

Reddit r/LocalLLaMA Models

Summary

A user shares their positive experience with low quantizations of Qwen 27B 3.8 on a Mac mini M4, using Unsloth's Q3 XXS quant, and asks for others' experiences with sub-Q3 quants.

For the GPU poor among us: I'm curious what results you're getting with low quants of Qwen 27B 3.8. My main inference hardware is limited (Mac mini M4 24 GB), but I'm getting great results with Unsloth's Q3 XXS. It's imperfect and makes minor mistakes, but it can work for hours autonomously towards a goal. And that's what really matters to me: A local LLM that I can trust to complete a goal. My context window size is about 180k. What are other people seeing? Is anyone getting anywhere with sub-Q3 quants?
Original Article

Similar Articles

Qwen3.6-27B Quantization Benchmark

Reddit r/LocalLLaMA

This article benchmarks various Qwen3.6-27B quantizations (Q8 to Q2) using KLD and Same Top P metrics, comparing providers like Unsloth and mradermacher, and offers recommendations for quality-size trade-offs.

Qwen 3.8 27b is strong even at Q3_xxs

Reddit r/LocalLLaMA

The user finds Qwen 3.8 27b in Q3 quantization highly effective for coding tasks with fast inference speeds, outperforming previous models, despite minor issues in general conversations.