Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark

Reddit r/LocalLLaMA Tools

Summary

mlx-dspark v0.10.0 adds support for Qwen3.8-27B on Apple Silicon, providing up to 3x faster inference through speculative decoding with lossless verification.

mlx-dspark is an MLX port of DeepSeek's DSpark speculative-decoding drafters (the DeepSpec release), plus z-lab's DFlash, with one lossless verify loop. v0.10.0 adds Qwen3.8-27B via RadixArk's drafter, the first SpecForge/SGLang-packaged head it loads. Numbers (M4 Pro 48 GB, medians of 3, greedy, output ids identical to plain decoding): 8-bit target: 2.45× mean at the auto-picked cap — 3.00× math / 2.38× code / 1.96× chat, 8.3 → 20.3 tok/s (code runs hit 3.18×). Peak ~29 GB. 4-bit target: 1.74× at 25.3 tok/s in ~18 GB (same drafter auto-resolves). Fun property: 8-bit + drafter (20-27 tok/s) beats plain 4-bit (14.6 tok/s) — 8-bit quality at better-than-4-bit speed. "Lossless" is checked, not asserted: the target verifies every drafted token, and the Mac app's Race view runs speculative vs plain on the same prompt and diffs the token ids (video is that view). Everything is pip install mlx-dspark (OpenAI-compatible server + Anthropic Messages API, so it can back Claude Code with a local model), and there's a native Mac app (DMG/Homebrew). Repo: github.com/ARahim3/mlx-dspark I'd appreciate any feedback you might have after using it.
Original Article

Similar Articles

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Reddit r/LocalLLaMA

Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.

harshatheg/Qwen-2.5-1B-RLCD

Hugging Face Models Trending

A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.