M5 ultra AI test results
Summary
M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.
Similar Articles
First M5 Ultra benchmarks
Unofficial benchmarks for the M5 Ultra chip show promising inference speeds with the Qwen 3.8 27B q4 model, achieving 50 tokens per second for threading and 1800 tokens per second for prefill at 8k context.
M5 Ultra and M6 Chip Benchmark Results Reveal Graphics Performance
Benchmark results for Apple's M5 Ultra and M6 chips show significant GPU performance gains, with the M5 Ultra achieving up to 59% higher Metal scores on Geekbench compared to the M3 Ultra, and the M6 chip scoring 34% higher than the M5.
The M5 Ultra Mac Studio tears through our benchmark tests
The M5 Ultra Mac Studio demonstrates exceptional benchmark performance, targeting AI developers and demanding workloads with its powerful specs and premium pricing.
M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents
The M5 Ultra Mac Studio is reviewed as a top-tier machine for running local AI agents, offering impressive performance and enabling cloud-free AI workflows. It is compared to the M3 Ultra and RTX 5090, with the reviewer praising its speed, thermal efficiency, and integration with AI models like Qwen3.8-Flash-Next.
@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.