@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …

X AI KOLs Timeline News

Summary

An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.

Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and there’s still more on the table. Running Qwen3.8-Flash-Next @ 4-bit locally. Now for the crazy bit: Artificial Analysis Intelligence Index v4.3: GPT-5.6 Sol 39 (Medium) GPT-5.6 Sol 42 (High) Qwen3.8-Flash-Next 42 GPT-5.6 Sol 44 (XHigh) GPT-5.6 Sol 47 (Max) Grok 4.5 39 (High) Claude Sonnet 5 38 (Max) *Qwen score currently estimated by Artificial Analysis. A 4-bit open-weights model running at 73 tok/s on a Mac under my desk... with an estimated intelligence score matching GPT-5.6 Sol at High reasoning. Local AI is getting ridiculous. 🔥 Plus, this model does not say no to a bit of red teaming. Once I am happy with it, I will do a pr for @jundotkim's oMLX :)
Original Article
View Cached Full Text

Cached at: 09/09/26, 03:48 AM

Been giving my M3 Ultra Studio some love and tinkering with MLX.

Started the afternoon at 42 tok/s. Now: 73 tok/s, and there’s still more on the table.

Running Qwen3.8-Flash-Next @ 4-bit locally.

Now for the crazy bit: Artificial Analysis Intelligence Index v4.3:

GPT-5.6 Sol 39 (Medium) GPT-5.6 Sol 42 (High) Qwen3.8-Flash-Next 42 GPT-5.6 Sol 44 (XHigh) GPT-5.6 Sol 47 (Max) Grok 4.5 39 (High) Claude Sonnet 5 38 (Max)

*Qwen score currently estimated by Artificial Analysis.

A 4-bit open-weights model running at 73 tok/s on a Mac under my desk… with an estimated intelligence score matching GPT-5.6 Sol at High reasoning.

Local AI is getting ridiculous. 🔥 Plus, this model does not say no to a bit of red teaming.

Once I am happy with it, I will do a pr for @jundotkim’s oMLX :)

Similar Articles

@nicekate8888: For the past twenty days, I've been obsessing over one thing — how to make Qwen3.6-27B run fast and well on my Mac. I started with Unsloth Q5, got 18 tok/s, and the fan was roaring. Then I switched to MLX 6bit + DFlash, hitting 22 tok/s, still not fast enough. Eventually I found MTPLX 4bit: 43 tok/s with good quality.

X AI KOLs Timeline

The user shares their experience optimizing Qwen3.6-27B inference speed on a Mac using different quantization methods (Unsloth Q5, MLX 6bit + DFlash, MTPLX 4bit), ultimately reaching 43 tok/s.