~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon

Reddit r/LocalLLaMA Tools

Summary

MTPLX framework achieves a 2x speed boost for Qwen3.8 27B and 1.5x for Qwen3.6 35B models on Apple Silicon, featuring auto-tuning and base conversion for MLX models.

https://x.com/koc_z3/status/2093581036756025744?s=46 ~ 2x speed boost for Qwen3.8 27B on Apple Silicon ~ 1.5x speed boost for Qwen3.6 35B AЗB Tested on an M1 Max 64GB Mac using MTPLX with 262K (MAX) Context length. Qwen3.8-27B (Q4): - Decode ~ 21 TPS - Prefill ~ 83 TPS (Peak 111 TPS) Qwen3.6-35B-A3B (Q4): - Decode ~ 55 TPS - Prefill ~ 300 TPS (Peak 623 TPS) Three key capabilities of this framework: Verified ~ 2x increase in local generation speed compared to base models. Auto-tuning: Determines the optimal MTP draft depth based on your specific chip, thermals, and memory bandwidth. Base Conversion: Transforms standard base models into MLX-ready MTP models. Repo: github.com/youssofal/MTPLX
Original Article

Similar Articles

@nash_su: Mac inference speed doubled. MTPLX is an integrated solution combining MLX and MTP, specifically optimized for model inference on Apple Silicon. By using models with a custom MTP head, it can deliver doubled inference speed. I tested it with Qwen3.6-27…

X AI KOLs Timeline

MTPLX is an integrated solution combining MLX and MTP, specifically optimized for model inference speed on Apple Silicon. Tests show that Qwen3.6-27B achieves double the inference speed of LM Studio, and it also integrates fan management.

Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit

Reddit r/LocalLLaMA

The user reviews a quantized and fine-tuned version of the Qwen3.6-35B model optimized for Apple Silicon via MLX, praising its speed, intelligence, and lack of safety disclaimers.