Tag
Jeff Geerling's benchmarks show Intel's Core 5 320 in a Dell XPS 13 matching or beating Apple Silicon on performance per watt, suggesting x86 can compete with ARM on efficiency.
Nativ is an open-source macOS app that runs Qwen 3.5 9B and other open models locally on Apple Silicon, offering customizable system prompts, telemetry, and integrations with coding agents, with no cloud or subscription required.
A developer updates MLX-Serve, a fast local inference server for Apple Silicon, to support recent models like LiquidAI 2.6B, MiniMax H3 video generation, and DeepSeek V4 Flash, with AntLing 3.0-flash coming soon.
Simon Willison highlights PipeNetwork/minimax-h3-mlx, a Python package that ports MiniMax-H3 to MLX for Apple Silicon, and demonstrates running it on an M5 Max MacBook Pro to generate a video from a text prompt.
A user shares experiments running a 4-bit quantized DeepSeek v4 Flash on a 32GB M5 MacBook Air, achieving roughly 50 tokens/s prefill and 1 token/s decode using streamed experts and other tricks.
Brian Roemmele reports that DeepSeek V4 Flash (304B, 1M context) now runs locally on Apple Silicon via the ds4 engine, sharing GGUF quantized builds with a fresh imatrix. The Hugging Face repo provides installation instructions and notes that these files are ds4-specific, not for llama.cpp.
TheTom releases an MLX quantized version of DeepSeek V4 Flash (284B MoE, 21B active) at 3.05 bpw, fitting in 101 GiB to run on 128GB Apple Silicon, with GGUF sibling also available.
mere.run is a local-first inference runtime for Apple Silicon and headless Linux that provides a single CLI for text, image, video, music, 3D, and more without requiring Python.
TurboFieldfare is an open-source Swift+Metal runtime that runs the Gemma 4 26B-A4B model on Apple Silicon Macs using only ~2GB of RAM by streaming experts from SSD, enabling inference on 8GB machines.
Qwen Scribe is an open-source tool for private, on-device transcription and dictation on Apple Silicon Macs, using Qwen3-ASR models via MLX. It supports drag-and-drop audio/video transcoding, language detection, SRT export, and system-wide dictation with a HUD.
A reported issue where virtual machines fail to boot when network mode is set to Bridged on Apple M5 Pro machines, affecting users of the UTM virtualization tool.
Rapid-MLX 0.11.0 brings major performance gains with prefix-cache and response caching, supports new model families including HY3 295B MoE and Qwen3-Coder-Next 80B, introduces structured output with guaranteed valid tool calls, and adds seamless integration with MCP servers for autonomous agent workflows.
Nativ is a free, open-source macOS app that lets you run frontier open AI models locally on Apple Silicon, with no accounts or subscriptions.
Nativ is a free, open-source macOS app that lets users run frontier open models locally on Apple Silicon, with no accounts, subscriptions, or cloud dependency.
A Twitter user highlights Todd Dailey, an ex-Apple engineer who recognized the potential of Apple Silicon for local AI and can now speak openly after leaving Apple.
elfuse is a lightweight, process-scoped runtime that runs Linux ELF binaries natively on macOS Apple Silicon without Docker or full VMs, using Hypervisor.framework and optionally Rosetta for x86_64 translation.
ntfsmac is an open-source tool that enables NTFS read/write on Apple Silicon macOS by using a libkrun microVM running ntfs-3g, avoiding kernel extensions and SIP modifications. It provides both CLI and a GUI menu-bar app.
A new model enables generating 3D models from a single image locally on Apple Silicon devices and iPhones, using less than 2GB RAM and completing in under 20 seconds.
Successfully ran the 75B Nemotron Puzzle model locally on a 64GB M2 Max Mac, demonstrating large model inference on consumer hardware.
Apple's Mac Studio desktop is now available with a 64GB unified memory configuration.