First M5 Ultra benchmarks

Reddit r/LocalLLaMA News

Summary

Unofficial benchmarks for the M5 Ultra chip show promising inference speeds with the Qwen 3.8 27B q4 model, achieving 50 tokens per second for threading and 1800 tokens per second for prefill at 8k context.

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!
Original Article

Similar Articles

M5 ultra AI test results

Reddit r/LocalLLaMA

M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.

GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra

Reddit r/LocalLLaMA

Optimizations for GLM 5.3 Flash on Apple M3 Ultra achieve up to 550 t/s prefill and 38 t/s inference speed through kernel fusion and efficient memory use, without quality loss.