M5 ultra AI test results

Reddit r/LocalLLaMA Products

Summary

M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.

So we have the results: - Prompt processing is up to 4 / 4.5 times faster than M3 ultra depending on how long the context is - tok/s is around 1.5x faster. But: the machine uses twice the power (400w vs 200w), makes more fan noise, and runs much hotter. https://www.youtube.com/watch?v=c_58D7ixOQI&t
Original Article

Similar Articles

First M5 Ultra benchmarks

Reddit r/LocalLLaMA

Unofficial benchmarks for the M5 Ultra chip show promising inference speeds with the Qwen 3.8 27B q4 model, achieving 50 tokens per second for threading and 1800 tokens per second for prefill at 8k context.

M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents

Hacker News Top

The M5 Ultra Mac Studio is reviewed as a top-tier machine for running local AI agents, offering impressive performance and enabling cloud-free AI workflows. It is compared to the M3 Ultra and RTX 5090, with the reviewer praising its speed, thermal efficiency, and integration with AI models like Qwen3.8-Flash-Next.