omlx

Tag

Cards List
#omlx

@ashxhart: @apple HomePod Mini has been liberated as a fun weekend project. Apple where kind enough to send me this HomePod Mini a…

X AI KOLs Timeline · 2026-09-01 Cached

A user modified a HomePod Mini to connect it to a local AI model GLM 5.3 Flash using a Spark cluster and oMLX for fluent AI conversations, mentioning that major AI companies wouldn't support this project.

0 favorites 0 likes
#omlx

Qwen3.8-27B thinking xhigh Vs. thinking off - Apple M5 Max

Reddit r/LocalLLaMA · 2026-08-29

Benchmarks on a MacBook Pro M5 Max show that disabling thinking mode in Qwen3.8-27B severely degrades output quality, while xhigh thinking mode uses 5.5x more tokens and runs 6x longer.

0 favorites 0 likes
#omlx

I asked Codex to optimize DeepSeek V4 Flash 8-bit MLX on oMLX. Got ~1.6x prefill and ~3x decode speedup.

Reddit r/LocalLLaMA · 2026-07-05

The author used Codex to optimize DeepSeek V4 Flash 8-bit MLX on oMLX, achieving approximately 1.6x prefill and 3x decode speedup.

0 favorites 0 likes
#omlx

@zhixianio: After receiving the new machine, I began an 'ascetic' practice of forcing myself to use local models for common tasks. I thought it would be painful, but both speed and quality greatly exceeded my expectations: Model: Qwen3.6-35B-A3B-oQ6-fp16-mtp, Running: oMLX, with N…

X AI KOLs Timeline · 2026-06-03 Cached

The author uses the Qwen3.6-35B-A3B model and oMLX tool on the new local machine for daily tasks, finding that both speed and quality far exceed expectations, even outperforming remote LLMs in PA and coding scenarios, demonstrating a significant improvement in on-device AI capabilities.

0 favorites 0 likes
#omlx

MLX engine comparison… and oMLX is the top choice.

Reddit r/LocalLLaMA · 2026-05-18

A blog post comparing MLX inference engines, concluding oMLX as the top choice, with benchmarks on M5 Max 64GB using Qwen3.6-35B-A3B-4bit.

0 favorites 0 likes
← Back to home

Submit Feedback