Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
Summary
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.
View Cached Full Text
Cached at: 09/12/26, 08:47 PM
Similar Articles
@nicebabycat: https://x.com/nicebabycat/status/2091726637155103126
This article provides a detailed test of the local deployment and performance of the Ling-3.0-tiny model on an Apple M5 chip Mac, demonstrating the feasibility of running a 7.9B parameter model at 47 tokens per second without a discrete GPU.
@ivanfioravanti: Apple M5 Max + MLX = raw power! Look at this demo I'm playing with "FasterLivePortrait-MLX" I started with MPS but resu…
The author demonstrates that migrating a LivePortrait implementation from MPS to Apple's MLX framework on an M5 Max chip results in significantly better performance and speed.
I benchmarked 21 local LLMs on a MacBook Air M5 for code quality AND speed
A developer benchmarked 21 local LLMs on MacBook Air M5 using HumanEval+ and found Qwen 3.6 35B-A3B (MoE) leads at 89.6% with 16.9 tok/s, while Qwen 2.5 Coder 7B offers the best RAM-to-performance ratio at 84.2% in 4.5 GB. Notably, Gemma 4 models significantly underperformed expectations (31.1% for 31B), possibly due to Q4_K_M quantization effects.
@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.
@Youssofal_: 72+ TPS on Qwen 3.6 27B on a Macbook pro M5 max. MTPLX V2 out now! The fastest way to run models on MLX.
MTPLX V2 is released, claiming 72+ tokens per second on Qwen 3.6 27B running on a Macbook Pro M5 Max via MLX.