M2 Ultra/Qwen3.8 Flash Next Update - latest oMLX introduces substantial speedup
Summary
The latest oMLX update introduces substantial speedups for Apple's M2 Ultra chip and the Qwen3.8 Flash AI model, enhancing performance and efficiency.
Similar Articles
Qwen3.8-Flash-Next optimised for Macs
The article details custom optimizations for running the Qwen3.8-Flash-Next AI model on Mac M1 Max hardware, including SSD streaming, custom quantizations, and a sparse attention mechanism to improve performance.
~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon
MTPLX framework achieves a 2x speed boost for Qwen3.8 27B and 1.5x for Qwen3.6 35B models on Apple Silicon, featuring auto-tuning and base conversion for MLX models.
M5 Ultra 96GB vs M5 Max 128GB — is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow?
The article compares Mac Studio M5 Ultra and M5 Max configurations for running AI models, focusing on trade-offs between bandwidth, RAM, and the upcoming Qwen3.8-Flash-Next model, with questions about performance and quantization.
@inco_ai: Hope you enjoy our first release! (and more to come)
Inco AI announces DFlash 2, a next-generation decoding technique that achieves up to 4.6× speed increase for Qwen3.8-27B on M5 Max MacBook Pro, improving efficiency without altering output.
Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070
The article details the merging of Multi-Token Prediction (MTP) support into ik_llama.cpp, which significantly enhances token generation speed for the Qwen3.8-Flash-Next model, with benchmarks showing up to double the speed on high-end hardware.