AI agents (mostly Opus 5.5) have been speeding up a 27B model from 66tok/s to 580tok/s on a Mac in 3 days by rewriting its inference engine

Reddit r/singularity News

Summary

AI agents, primarily using Opus 5.5, have significantly accelerated a 27B model's inference speed from 66 to 580 tokens per second on a Mac by rewriting its inference engine in just three days.

https://www.yukon.org/mlxfast
Original Article

Similar Articles

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Reddit r/LocalLLaMA

Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.