@RayFernando1337: Let him cook. This is a fun time to be owning 2 DGX Sparks RN
Summary
A Twitter user discusses owning two NVIDIA DGX Sparks systems, while another user reports improved token throughput with a model update to DFlash2-7, achieving 67 tokens per second.
View Cached Full Text
Cached at: 08/30/26, 12:00 AM
Let him cook. This is a fun time to be owning 2 DGX Sparks RN
Sufyan (@sfxnz): Default is now DFlash2-7 at two sequences.
Just measured 67 tok/s at c=1 structured. Was 51 tok/s on DFlash2-5 at four sequences.
Staying at 327k context.
Enjoy.
Similar Articles
@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …
MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.
@TechMDAI: The DGX spark community is on Fire
The DGX spark community is experiencing high activity and enthusiasm.
1 rtx pro 6000 or 2 dgx sparks
A comparison between a single RTX Pro 6000 GPU and two DGX Spark systems for AI compute tasks.
@RayFernando1337: I’m super excited about Local Studio! I have two DGX Sparks connected together, and I will be able to run my own agents…
Ray Fernando expresses excitement about Local Studio, an open-source interface for running AI agents on two connected DGX Sparks, similar to Codex.
@MiaAI_lab: What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s…
A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.