@AlmustyFX: This is the kind of local AI test that actually matters. A 7.9B MoE running at 152 tok/s on an M4 Pro with 64GB unified…

X AI KOLs Timeline News

Summary

A 7.9B MoE model runs at 152 tokens per second on an M4 Pro with 64GB unified memory, enabling offline processing of sensitive contract data and demonstrating the practical use of local AI.

This is the kind of local AI test that actually matters. A 7.9B MoE running at 152 tok/s on an M4 Pro with 64GB unified memory, while handling sensitive contract data completely offline. No cloud. No data leaving the machine. Just a lightweight model doing useful work in under a second. Local AI is becoming less of a demo and more of a real tool.
Original Article
View Cached Full Text

Cached at: 08/29/26, 04:07 PM

This is the kind of local AI test that actually matters.

A 7.9B MoE running at 152 tok/s on an M4 Pro with 64GB unified memory, while handling sensitive contract data completely offline.

No cloud. No data leaving the machine. Just a lightweight model doing useful work in under a second.

Local AI is becoming less of a demo and more of a real tool.

洛一可🍥 (@luoyike2003): 手痒把 Ling-3.0-tiny 的 MLX 4bit 拉到本机跑了一圈,把流程、选型理由和踩的坑一起记下来

【机器】Apple M4 Pro,64G 统一内存,macOS 26.5.1

【为什么挑 4bit 这一版】

Ling-3.0-tiny 是 7.9B 总参、每 token 只激活 1.3B 的稀疏 MoE:128 个路由专家取 top-8 加 1 个共享专家,24 层,注意力是

Similar Articles