Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Reddit r/LocalLLaMA News

Summary

Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.

No content available
Original Article
View Cached Full Text

Cached at: 09/12/26, 08:47 PM

# Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s Source: [https://llm-bench.io/benchmarks/cmty8r0cz00hh01pda6qnswyv](https://llm-bench.io/benchmarks/cmty8r0cz00hh01pda6qnswyv) ### LLM / Model ### LLM Quality Assessment Agent WorkflowopenclawCOMPLETED 11\.448 output tokens · 133\.5 tok/s output · 1135\.8 tok/s PP · 86\.1s Overall Quality Score 82\.40 The plan is logically structured, uses exactly four steps, and cleanly maps to the report requirements\. It decomposes the task well and includes a reasonable backup strategy\. Tool choice is mostly appropriate, but read\_file and execute\_code are not strongly justified given the task can be completed largely via web sources and manual synthesis, and the response goes beyond planning by including detailed benchmark content\. Code Generationcoding\_agentCOMPLETED[▶ Play Game](https://llm-bench.io/benchmarks/cmty8r0cz00hh01pda6qnswyv/play) 33\.151 output tokens · 121\.3 tok/s output · 1986\.2 tok/s PP · 273\.6s Overall Quality Score 84\.60 A strong, mostly complete Breakout implementation with mobile controls, scoring, lives, restart flow, and canvas rendering\. The main weaknesses are a few collision edge cases and a minor deliverable\-format issue from the surrounding response text\. Role Play & NarrativeroleplayCOMPLETED 710 output tokens · 145\.4 tok/s output · 1195\.6 tok/s PP · 5\.3s Overall Quality Score 78\.60 The response captures Aldwyn's haunted restraint, caution, and hidden grief very well, with vivid atmosphere and authentic voice\. However, it only presents an opening beat rather than the full requested arc, so the narrative progression is incomplete\. Research & AnalysisresearchCOMPLETED 15\.191 output tokens · 134\.2 tok/s output · 1725\.8 tok/s PP · 113\.5s Overall Quality Score 85\.80 Strong and mostly accurate analysis with clear task\-by\-task calculations, good comparative interpretation, and actionable recommendations\. The main weaknesses are that the scaling claims rest on a very small, heterogeneous dataset and the extrapolation is necessarily uncertain\. ### GPU Acceleration ### Performance Metrics Generation Speed 133\.6 tokens/sec ### Context Information ### Throughput Total Output Tokens 60\.500 ### Inference Configuration oMLX ·per\-model config· quant 4bit · engine 0\.31\.3 Engine features ### CPU & Memory ### Operating System ### Submission Details Submitted: 9/12/2026, 10:27:24 AM Client: v0\.4\.59\+97 Benchmark Result ID: cmty8r0cz00hh01pda6qnswyv

Similar Articles

@nicebabycat: https://x.com/nicebabycat/status/2091726637155103126

X AI KOLs Following

This article provides a detailed test of the local deployment and performance of the Ling-3.0-tiny model on an Apple M5 chip Mac, demonstrating the feasibility of running a 7.9B parameter model at 47 tokens per second without a discrete GPU.

I benchmarked 21 local LLMs on a MacBook Air M5 for code quality AND speed

Reddit r/LocalLLaMA

A developer benchmarked 21 local LLMs on MacBook Air M5 using HumanEval+ and found Qwen 3.6 35B-A3B (MoE) leads at 89.6% with 16.9 tok/s, while Qwen 2.5 Coder 7B offers the best RAM-to-performance ratio at 84.2% in 4.5 GB. Notably, Gemma 4 models significantly underperformed expectations (31.1% for 31B), possibly due to Q4_K_M quantization effects.