@YRSM_Simon: Crazy
Summary
DeepSeek-V4-Flash-DSpark achieves 328 tok/s single inference and 1.7k tok/s batch throughput on 4x RTX PRO 6000 GPUs.
View Cached Full Text
Cached at: 07/09/26, 09:40 AM
Crazy
AI✖️Satoshi⏩️ (@AiXsatoshi): DeepSeek-V4-Flash-DSpark on 4x RTX PRO 6000 has come quite far.
- Single inference: ~328 tok/s
- Batch throughput: ~1.7k tok/s
Similar Articles
@YRSM_Simon: 120 t/s ! Good job, @UnslothAI
Unsloth AI announces DSpark, enabling DeepSeek-V4-Flash GGUF models to run ~1.4–2× faster locally, reaching 120 tokens/s with no accuracy change.
@Snixtp: DeepSeek V4 Flash on a single RTX Pro 6000?
DeepSeek V4 Flash GGUF quantizations have been released by antirez, enabling the model to run on single GPUs like the RTX Pro 6000 and Macs with 128GB+ RAM. The quantized files are available on Hugging Face with instructions for the DS4 inference engine.
@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …
MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.
@SuJinYan123: Just 6 hours after DeepSeek open-sourced the Qwen DSpark weights, OpenInfer already has DSpark support running on RTX 5…
OpenInfer, a pure Rust+CUDA LLM inference engine, quickly added support for DeepSeek's DSpark speculative decoding technique on RTX 5090, achieving nearly 500 tok/s per user and scaling to ~2.4K aggregate tok/s, outperforming DFlash on non-random workloads.
@no_stp_on_snek: It's still pretty awesome what a single spark can do. Thanks @NVIDIAAI
A user shares that they replicated running DeepSeek-V4-Flash-0731 on a DGX Spark using antirez's DwarfStar-4 setup, confirming impressive performance on a single device.