Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark
Summary
mlx-dspark v0.10.0 adds support for Qwen3.8-27B on Apple Silicon, providing up to 3x faster inference through speculative decoding with lossless verification.
Similar Articles
~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon
MTPLX framework achieves a 2x speed boost for Qwen3.8 27B and 1.5x for Qwen3.6 35B models on Apple Silicon, featuring auto-tuning and base conversion for MLX models.
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro
Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.
harshatheg/Qwen-2.5-1B-RLCD
A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.
@LinusEkenstam: Quite a big deal. 1.35x faster inference than MLX-LM 1.23x faster at prefill having the hybrid option to pick from insi…
Perplexity has open-sourced Lily, a local inference engine optimized for Qwen3.6-35B-A3B on Apple Silicon, achieving 1.35x faster inference than MLX-LM.
Perplexity open-sourced their Mac inference server for Qwen 3.6
Perplexity has open-sourced a Mac inference server optimized for the Qwen 3.6 model to achieve best performance on Apple Silicon.