You can now convert EXL3 quants on Apple Silicon Mac
Summary
A new tool enables converting and running EXL3 quantized models on Apple Silicon Macs, matching or nearly matching RTX conversion quality, making high-fidelity quants more accessible.
Similar Articles
I ported EXL3 to run well on Apple Silicon - PonyExl3
Ported the EXL3 LLM codec to run on Apple Silicon via Metal, achieving high prefill and generation speeds on M5 Max (e.g., ~600 tok/s prefill, 17-80 tok/s gen on various models).
Run Qwen3.8 27B locally: real numbers from my Mac Studio
The article provides real-world performance benchmarks for running the Qwen3.8 27B AI model locally on a Mac Studio, comparing it to its predecessor and discussing hardware requirements and quantization effects.
~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon
MTPLX framework achieves a 2x speed boost for Qwen3.8 27B and 1.5x for Qwen3.6 35B models on Apple Silicon, featuring auto-tuning and base conversion for MLX models.
MAX models can now run on Apple silicon GPUs
MAX models have been updated to run on Apple silicon GPUs, enabling faster inference on Macs.
Perplexity open-sourced their Mac inference server for Qwen 3.6
Perplexity has open-sourced a Mac inference server optimized for the Qwen 3.6 model to achieve best performance on Apple Silicon.