native-mtp

Tag

Cards List
#native-mtp

@zhixianio: After receiving the new machine, I began an 'ascetic' practice of forcing myself to use local models for common tasks. I thought it would be painful, but both speed and quality greatly exceeded my expectations: Model: Qwen3.6-35B-A3B-oQ6-fp16-mtp, Running: oMLX, with N…

X AI KOLs Timeline · 2026-06-03 Cached

The author uses the Qwen3.6-35B-A3B model and oMLX tool on the new local machine for daily tasks, finding that both speed and quality far exceed expectations, even outperforming remote LLMs in PA and coding scenarios, demonstrating a significant improvement in on-device AI capabilities.

0 favorites 0 likes
#native-mtp

I added native MTP to exo for Qwen3.6 MLX models; here are the exactness and speed results

Reddit r/LocalLLaMA · 2026-05-23

Added native multi-token prediction (MTP) support to the exo local inference tool for Qwen3.6 MLX models, achieving up to 2x speedup on 27B models on an M5 Max laptop while maintaining exactness.

0 favorites 0 likes
← Back to home

Submit Feedback