@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …
Summary
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.
View Cached Full Text
Cached at: 09/09/26, 03:48 AM
Been giving my M3 Ultra Studio some love and tinkering with MLX.
Started the afternoon at 42 tok/s. Now: 73 tok/s, and there’s still more on the table.
Running Qwen3.8-Flash-Next @ 4-bit locally.
Now for the crazy bit: Artificial Analysis Intelligence Index v4.3:
GPT-5.6 Sol 39 (Medium) GPT-5.6 Sol 42 (High) Qwen3.8-Flash-Next 42 GPT-5.6 Sol 44 (XHigh) GPT-5.6 Sol 47 (Max) Grok 4.5 39 (High) Claude Sonnet 5 38 (Max)
*Qwen score currently estimated by Artificial Analysis.
A 4-bit open-weights model running at 73 tok/s on a Mac under my desk… with an estimated intelligence score matching GPT-5.6 Sol at High reasoning.
Local AI is getting ridiculous. 🔥 Plus, this model does not say no to a bit of red teaming.
Once I am happy with it, I will do a pr for @jundotkim’s oMLX :)
Similar Articles
@Prince_Canuma: My home compute for MLX and research: • M3 Ultra — 512GB (sponsored by community + @wai_protocol) • RTX PRO 6000 — 96GB…
A researcher shares their home compute setup for MLX and AI research, featuring M3 Ultra with 512GB, RTX PRO 6000 with 96GB, and M3 Max with 96GB for model porting and stress testing.
M2 Ultra/Qwen3.8 Flash Next Update - latest oMLX introduces substantial speedup
The latest oMLX update introduces substantial speedups for Apple's M2 Ultra chip and the Qwen3.8 Flash AI model, enhancing performance and efficiency.
Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.
@rohanpaul_ai: Qwen 3.6 27B on a MacBook Pro M5 Max 64GB hitting 34tokens per sec, locally with atomic[.]chat 90% acceptance rate, i.e…
Qwen 3.6 27B achieves 34 tokens/sec on a MacBook Pro M5 Max 64GB locally with 90% draft acceptance, enabled by TurboQuant, GGUF, and llama.cpp, showcasing a major advancement in laptop-based AI inference.
@nicekate8888: For the past twenty days, I've been obsessing over one thing — how to make Qwen3.6-27B run fast and well on my Mac. I started with Unsloth Q5, got 18 tok/s, and the fan was roaring. Then I switched to MLX 6bit + DFlash, hitting 22 tok/s, still not fast enough. Eventually I found MTPLX 4bit: 43 tok/s with good quality.
The user shares their experience optimizing Qwen3.6-27B inference speed on a Mac using different quantization methods (Unsloth Q5, MLX 6bit + DFlash, MTPLX 4bit), ultimately reaching 43 tok/s.