QWEN3.6-27B-MLX-8bit (29.5GIG) Is Excellent

Reddit r/openclaw Models

Summary

A user shares their experience switching to the QWEN3.6-27B-MLX-8bit model for local AI tasks, finding it performs comparably to larger models while saving significant RAM on a Mac Studio, improving workflow stability.

For anyone using local models for orchestration, research, and writing, I had been using QWEN3.5-122B-a10-uncensored-hauhaucs-aggressive (79GIG)as my main driver. It's an excellent model and drives eight hour workflows reliably. With an M3Ultra Mac Studio with 256 GIG of RAM, however, when I'd get to the video generation portion of my workflow, I would easily push 93% of memory usage and periodically be forced to start closing applications. I did a lot of testing the past few days (I posted about it recently) and found no other models amongst QWEN or Nemotron that were comparable. However, today I found that QWEN3.6-27B-MLX-8big (29.5GIG) performed as good at all of these tasks, but was slower. However, until I get my newer Mac Studio to add to this one dedicated only to video work, it gives me back 50 GIG of RAM to work with. That frees up a lot of headspace to avoid crashes. Just sharing, for those who embrace local models.
Original Article

Similar Articles

Run Qwen3.8 27B locally: real numbers from my Mac Studio

Hacker News Top

The article provides real-world performance benchmarks for running the Qwen3.8 27B AI model locally on a Mac Studio, comparing it to its predecessor and discussing hardware requirements and quantization effects.

Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit

Reddit r/LocalLLaMA

The user reviews a quantized and fine-tuned version of the Qwen3.6-35B model optimized for Apple Silicon via MLX, praising its speed, intelligence, and lack of safety disclaimers.