@MiaAI_lab: What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s…
Summary
A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.
View Cached Full Text
Cached at: 06/28/26, 03:59 AM
What’s the best model you can run on your @NVIDIAAI DGX Spark?
1× DGX Spark
- Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s
- DeepSeek v4 Flash REAP
- Qwen 3.6 27b - 256k ctx 19, tok/s
2× DGX Sparks ← sweet spot!
- DeepSeek v4 Flash - 1M ctx, 40-45 tok/s
- Step-3.7-Flash - 256k ctx, image support, 30 tok/s
4× DGX Sparks
- GLM 5.2 NVFP4 on all 4
- 2× DeepSeek v4 Flash setups, each on 2× DGX Sparks.
Links and repos below
Similar Articles
dgx sparks and new models my tests and results
This article presents test results for AI models like DeepSeek V4 Flash and Qwen3.8 on NVIDIA DGX Sparks hardware, detailing performance metrics, context lengths, and benchmark scores with operational insights.
@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …
MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.
@MiaAI_lab: Nvidia did it again! @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on D…
Nvidia's optimized Qwen 3.6 27B NVFP4 model achieves 41% faster single-session inference and 23-25% faster concurrent inference on DGX Spark compared to Unsloth's version.
@ivanfioravanti: DGX Spark Context Benchmark on Qwen3.6-35B-A3B-UD-Q8_K_XL llamacpp script released by Mia. It's fast! Time to test qual…
Benchmark results for Qwen3.6-35B-A3B-UD-Q8_K_XL on DGX Spark using llama.cpp script by Mia, showing fast token generation times across various context lengths.
DGX Spark agentic usage numbers
A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.