MLX engine comparison… and oMLX is the top choice.
Summary
A blog post comparing MLX inference engines, concluding oMLX as the top choice, with benchmarks on M5 Max 64GB using Qwen3.6-35B-A3B-4bit.
Similar Articles
oMLX
oMLX is an open-source Mac application that acts as an LLM inference server, significantly reducing response times for AI coding tools from 90s to about 5s using a RAM+SSD tiered KV cache.
@AlexJonesax: Qwen3.6-27b absolutely flying on a M5Max with MTP enabled & oMLX inference.
A community report highlights high inference performance for the Qwen3.6-27b model on M5Max hardware using oMLX optimization.
@AlexJonesax: Two open-source MLX inference servers worth knowing about if you run LLMs on Mac: MTPLX (@youssofal) Uses a model's own…
This article highlights two open-source MLX inference servers for Mac: MTPLX, which optimizes token speed using speculative decoding without a draft model, and oMLX, which improves workflow efficiency with persistent KV caches for coding agents.
@jun_song: The new engine for MLX is in its final stages of development. Just ran GLM-5.2 on a single MacBook (116GB) hitting 41.8…
Jun Song announces the final development stage of a new MLX engine, achieving 41.8 tok/s on a MacBook with a 256k context window and only ~4% quality loss, representing a significant performance improvement.
@LinusEkenstam: Quite a big deal. 1.35x faster inference than MLX-LM 1.23x faster at prefill having the hybrid option to pick from insi…
Perplexity has open-sourced Lily, a local inference engine optimized for Qwen3.6-35B-A3B on Apple Silicon, achieving 1.35x faster inference than MLX-LM.