@svpino: DeepSeek-V4-Flash running at 5.71 token/s on a Mac M5 Pro. Every day, we get better models running on consumer hardware…

X AI KOLs Timeline News

Summary

The tweet highlights DeepSeek-V4-Flash running at 5.71 tokens per second on a Mac M5 Pro, emphasizing advancements in local AI inference on consumer hardware, with a mention of Tencent's open-source Palm-Infra for Apple Silicon optimization.

DeepSeek-V4-Flash running at 5.71 token/s on a Mac M5 Pro. Every day, we get better models running on consumer hardware. And we aren't talking about toy models anymore: we can already run a ton of intelligence without paying for an external GPU.
Original Article
View Cached Full Text

Cached at: 08/25/26, 02:02 AM

DeepSeek-V4-Flash running at 5.71 token/s on a Mac M5 Pro.

Every day, we get better models running on consumer hardware.

And we aren’t talking about toy models anymore: we can already run a ton of intelligence without paying for an external GPU.

Tencent Cloud (@tencentcloud): From dense models to 122B and 295B MoE, local inference should work with the hardware you have—not require a datacenter. Palm-Infra is open source. Built by Tencent’s #PalmAI team, it brings Metal, CPU SIMD and SSD expert offload to Apple Silicon.

Similar Articles

DeepSeek-V4-Flash 284B on 5.3GB of memory

Reddit r/LocalLLaMA

A developer showcases Mference, a new inference engine that runs MoE models like DeepSeek-V4-Flash on just ~5.3GB of memory by streaming experts from SSD, with a native Mac app and OpenAI-compatible server.

@danveloper: https://x.com/danveloper/status/2064387956387758206

X AI KOLs Timeline

A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.