@MiaAI_lab: What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s…
Summary
A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.
View Cached Full Text
Cached at: 06/28/26, 03:59 AM
What’s the best model you can run on your @NVIDIAAI DGX Spark?
1× DGX Spark
- Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s
- DeepSeek v4 Flash REAP
- Qwen 3.6 27b - 256k ctx 19, tok/s
2× DGX Sparks ← sweet spot!
- DeepSeek v4 Flash - 1M ctx, 40-45 tok/s
- Step-3.7-Flash - 256k ctx, image support, 30 tok/s
4× DGX Sparks
- GLM 5.2 NVFP4 on all 4
- 2× DeepSeek v4 Flash setups, each on 2× DGX Sparks.
Links and repos below
Similar Articles
@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …
MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.
@MiaAI_lab: Nvidia did it again! @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on D…
Nvidia's optimized Qwen 3.6 27B NVFP4 model achieves 41% faster single-session inference and 23-25% faster concurrent inference on DGX Spark compared to Unsloth's version.
@ivanfioravanti: DGX Spark Context Benchmark on Qwen3.6-35B-A3B-UD-Q8_K_XL llamacpp script released by Mia. It's fast! Time to test qual…
Benchmark results for Qwen3.6-35B-A3B-UD-Q8_K_XL on DGX Spark using llama.cpp script by Mia, showing fast token generation times across various context lengths.
DGX Spark agentic usage numbers
A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.
@TheAhmadOsman: Dense models like Qwen 3.8 27B are a TERRIBLE experience on unified-memory systems like DGX Spark btw DGX Sparks are be…
Ahmad Osman argues that dense models like Qwen 27B perform poorly on unified-memory systems such as NVIDIA DGX Spark, suggesting MoE models are a better fit; he claims discrete GPUs like the RTX PRO 6000 deliver far better performance for agentic workloads.