mlx

Tag

Cards List
#mlx

@Nativ_AI: Run Qwen 3.5 9B locally with fully customizable system prompts. Make it sing. Make it rhyme. Make it brainstorm, write,…

X AI KOLs Timeline · yesterday Cached

Nativ is an open-source macOS app that runs Qwen 3.5 9B and other open models locally on Apple Silicon, offering customizable system prompts, telemetry, and integrations with coding agents, with no cloud or subscription required.

0 favorites 0 likes
#mlx

nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx

Reddit r/LocalLLaMA · 2d ago

The author wrote a pure-MLX runtime for the vision and audio towers of Nvidia's Nemotron Omni model, enabling the full multimodal model to run locally on Apple Silicon Macs. All 23 tests pass against the PyTorch reference, and it achieves 67-152 tok/s on an M5 Max.

0 favorites 0 likes
#mlx

@ddalcu: What a crazy last 4 days for local AI... insane... https://github.com/ddalcu/mlx-serve/releases/tag/v26.8.2… @liquidai …

X AI KOLs Timeline · 4d ago Cached

A developer updates MLX-Serve, a fast local inference server for Apple Silicon, to support recent models like LiquidAI 2.6B, MiniMax H3 video generation, and DeepSeek V4 Flash, with AntLing 3.0-flash coming soon.

0 favorites 0 likes
#mlx

PipeNetwork/minimax-h3-mlx

Simon Willison's Blog · 4d ago Cached

Simon Willison highlights PipeNetwork/minimax-h3-mlx, a Python package that ports MiniMax-H3 to MLX for Apple Silicon, and demonstrates running it on an M5 Max MacBook Pro to generate a video from a text prompt.

0 favorites 0 likes
#mlx

@no_stp_on_snek: https://huggingface.co/thetom-ai/DeepSeek-V4-Flash-ConfigI-MLX… Fyi. Fits on 128GB of ram for metal. GGUF is coming jus…

X AI KOLs Following · 2026-08-01 Cached

TheTom releases an MLX quantized version of DeepSeek V4 Flash (284B MoE, 21B active) at 3.05 bpw, fitting in 101 GiB to run on 128GB Apple Silicon, with GGUF sibling also available.

0 favorites 0 likes
#mlx

Show HN: Qwen Scribe – local transcription and dictation for Apple Silicon

Hacker News Top · 2026-07-29 Cached

Qwen Scribe is an open-source tool for private, on-device transcription and dictation on Apple Silicon Macs, using Qwen3-ASR models via MLX. It supports drag-and-drop audio/video transcoding, language detection, SRT export, and system-wide dictation with a HUD.

0 favorites 0 likes
#mlx

@eisokant: Excited to launch http://MLX.fast with @eigenlabs today. It's an open autoresearch competition to make Laguna XS 2.1 in…

X AI KOLs Following · 2026-07-28 Cached

Eigen Labs launches an open autoresearch competition called MLX.fast to optimize inference speed of the Laguna XS 2.1 model on consumer Macs, aiming to make it as fast as possible via community contributions.

0 favorites 0 likes
#mlx

@Raullen: Rapid-MLX 0.11.0 is out! Making local models on Apple Silicon reliable enough to run your agent workflows, not just dem…

X AI KOLs Following · 2026-07-24 Cached

Rapid-MLX 0.11.0 brings major performance gains with prefix-cache and response caching, supports new model families including HY3 295B MoE and Qwen3-Coder-Next 80B, introduces structured output with guaranteed valid tool calls, and adds seamless integration with MCP servers for autonomous agent workflows.

0 favorites 0 likes
#mlx

Apple M5 isn't making full use of its matmul cores yet

Reddit r/LocalLLaMA · 2026-07-23

Apple M5 silicon supports INT8 activations for matrix multiplication, but inference backends like MLX and Llama.cpp currently use 16-bit; custom w8a8 kernels achieve up to 1.4x speedup on Gemma4 prefill tasks.

0 favorites 0 likes
#mlx

@jundotkim: oMLX 0.5.2 is out. (Sorry for the long silence!) https://github.com/jundot/omlx/releases… oMLX is the most convenient w…

X AI KOLs Timeline · 2026-07-21 Cached

oMLX 0.5.2 release adds live menu bar activity, a reorganized Models menu, Bonsai low-bit kernels, and improved performance with custom Metal kernels and native speculative decoding, making it the fastest way to run MLX models on Mac.

0 favorites 0 likes
#mlx

Nativ: Run AI models locally on your Mac

Simon Willison's Blog · 2026-07-21 Cached

Nativ is a new macOS desktop app that wraps MLX to run AI models locally, offering a chat interface and API server.

0 favorites 0 likes
#mlx

Models to get before any political disruption

Reddit r/LocalLLaMA · 2026-07-21

A curated list of recommended AI models (e.g., Qwen3.5, Gemma 4, Mistral Medium) for systems with up to 256GB unified RAM, covering creative, coding, agentic, and general chat use cases.

0 favorites 0 likes
#mlx

Nativ: Run frontier open models locally on your Mac

Hacker News Top · 2026-07-20 Cached

Nativ is a free, open-source macOS app that lets you run frontier open AI models locally on Apple Silicon, with no accounts or subscriptions.

0 favorites 0 likes
#mlx

@Prince_Canuma: Excited to introduce Nativ Run frontier open models locally on your Mac. No accounts, no subscriptions, no cloud. Built…

X AI KOLs Timeline · 2026-07-20 Cached

Nativ is a free, open-source macOS app that lets users run frontier open models locally on Apple Silicon, with no accounts, subscriptions, or cloud dependency.

0 favorites 0 likes
#mlx

@no_stp_on_snek: anyone still talking about mlx-swift-lm? said i was taking the day off... cleaned the chicken coop, got a workout in, f…

X AI KOLs Following · 2026-07-16 Cached

The author describes implementing TurboQuant KV-cache compression into Apple's mlx-swift-lm, achieving 2.7x compression with quality on par with 8-bit, and 3-4x decode speed improvements via a fused Metal kernel.

0 favorites 0 likes
#mlx

Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable

Reddit r/LocalLLaMA · 2026-07-13

Fixed three bugs in a qMLX fork for running Qwen3.5-122B on Mac Studio, reducing prefill time from minutes to sub-seconds for long-context inference; open-sourced the fork and benchmark script.

0 favorites 0 likes
#mlx

@jun_song: The new engine for MLX is in its final stages of development. Just ran GLM-5.2 on a single MacBook (116GB) hitting 41.8…

X AI KOLs Following · 2026-07-09 Cached

Jun Song announces the final development stage of a new MLX engine, achieving 41.8 tok/s on a MacBook with a 256k context window and only ~4% quality loss, representing a significant performance improvement.

0 favorites 0 likes
#mlx

@_ARahim_: DeepSeek's DSpark speculative-decoding drafters, benchmarked on a Mac Native on Apple Silicon (MLX), lossless; identica…

X AI KOLs Timeline · 2026-07-07 Cached

mlx-dspark brings DeepSeek's DSpark and z-lab's DFlash speculative decoding drafters to Apple Silicon via MLX, enabling lossless speedup (~1.4–1.6×, up to 2× on code/math) and an OpenAI-compatible API for local inference.

0 favorites 0 likes
#mlx

@Youssofal_: 72+ TPS on Qwen 3.6 27B on a Macbook pro M5 max. MTPLX V2 out now! The fastest way to run models on MLX.

X AI KOLs Following · 2026-07-07 Cached

MTPLX V2 is released, claiming 72+ tokens per second on Qwen 3.6 27B running on a Macbook Pro M5 Max via MLX.

0 favorites 0 likes
#mlx

@Prince_Canuma: mlx-vlm v0.6.4 is here! Get started today: > uv pip install -U mlx-vlm 5 new model families — MiniMax M3, Kimi K2.5, Un…

X AI KOLs Timeline · 2026-07-06 Cached

mlx-vlm v0.6.4 is released with support for 5 new model families, TTS/STT endpoints, and significant performance improvements including TurboQuant and continuous batching.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback