@VincentLogic: Discovered an amazing open-source project! Redis creator antirez made a splash! ds4 — DeepSeek V4 Flash local inference engine, optimized for Mac Metal, topping GitHub charts for days! And here's the killer part: 128GB…

X AI KOLs Timeline Tools

Summary

Redis creator antirez released an open-source project called ds4, a DeepSeek V4 Flash local inference engine optimized for Mac Metal, featuring disk KV caching, ultra-long context, and excellent performance.

Discovered an amazing open-source project! Redis creator antirez made a splash! ds4 — DeepSeek V4 Flash local inference engine, optimized for Mac Metal, topping GitHub charts for days! And here's the killer part: A 128GB MacBook Pro M3 Max can run the full DeepSeek V4 at 26 tokens/s generation speed, 58 t/s prefill, supporting 1 million token ultra-long context! The core technology is incredibly hardcore: ✦ Disk KV cache: Uses SSD as VRAM, breaking memory limits ✦ Written in pure C: Extreme performance optimization ✦ Compatible with OpenAI API: Seamless integration with existing Agent tools ✦ Highly compressed KV cache: Memory footprint is extremely stable Even with a 1 million token context, inference is smooth and the memory footprint is rock stable. Mac users who want to experience the pinnacle of local AI must give this a try! Project link in the comments.
Original Article

Similar Articles

@PandaTalk8: These test results are stunning. The original poster tested the DS4 inference engine written in C by @antirez, and local deployment seems incredibly fast. The good news is that only 128GB of RAM is needed to run a local model equivalent to GPT-4o. The bad news is that you need a MacBook Pro with 128GB of RAM.

X AI KOLs Timeline

This article reports on tests of the DS4 inference engine written in C by @antirez, noting its impressive speed when running a GPT-4o-equivalent model on a MacBook Pro with 128GB of RAM.