First M5 Ultra benchmarks
Summary
Unofficial benchmarks for the M5 Ultra chip show promising inference speeds with the Qwen 3.8 27B q4 model, achieving 50 tokens per second for threading and 1800 tokens per second for prefill at 8k context.
Similar Articles
M5 ultra AI test results
M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.
M5 Ultra and M6 Chip Benchmark Results Reveal Graphics Performance
Benchmark results for Apple's M5 Ultra and M6 chips show significant GPU performance gains, with the M5 Ultra achieving up to 59% higher Metal scores on Geekbench compared to the M3 Ultra, and the M6 chip scoring 34% higher than the M5.
Qwen3.8-Next streaming - 150tps prefill, 3.6 tps decode on M5 Air
A user tested the Qwen3.8-Next model on an Apple M5 Air with 3-bit quantization, achieving 150 tokens per second prefill and 3.6 tps decode, outperforming a dense 27b model in some metrics.
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
Optimizations for GLM 5.3 Flash on Apple M3 Ultra achieve up to 550 t/s prefill and 38 t/s inference speed through kernel fusion and efficient memory use, without quality loss.
Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.