MLX engine comparison… and oMLX is the top choice.

Reddit r/LocalLLaMA Tools

Summary

A blog post comparing MLX inference engines, concluding oMLX as the top choice, with benchmarks on M5 Max 64GB using Qwen3.6-35B-A3B-4bit.

Just stumbled on this blog. A very interesting read if you are picking inference engine. M5 Max 64GB with mlx-community/Qwen3.6-35B-A3B-4bit. The MTPLX in the article use 3.6 27B so it's not apple to apple. https://preview.redd.it/huxhasc4gx1h1.png?width=990&format=png&auto=webp&s=88cf7828b18eb8dea7a4c92c041f2b5c795f1824 https://preview.redd.it/fhevre6agx1h1.png?width=990&format=png&auto=webp&s=7bbc9aecbb5684aeeedf712e5a1017d0aab68fa7 [https://www.largitdata.com/blog\_detail/20260511](https://www.largitdata.com/blog_detail/20260511)
Original Article

Similar Articles

oMLX

Product Hunt

oMLX is an open-source Mac application that acts as an LLM inference server, significantly reducing response times for AI coding tools from 90s to about 5s using a RAM+SSD tiered KV cache.