@MiaAI_lab: What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * ⁠Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s…

X AI KOLs Timeline News

Summary

A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.

What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * ⁠Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s * DeepSeek v4 Flash REAP * ⁠Qwen 3.6 27b - 256k ctx 19, tok/s 2× DGX Sparks ← sweet spot! * DeepSeek v4 Flash - 1M ctx, 40-45 tok/s * Step-3.7-Flash - 256k ctx, image support, 30 tok/s 4× DGX Sparks * GLM 5.2 NVFP4 on all 4 * 2× DeepSeek v4 Flash setups, each on 2× DGX Sparks. Links and repos below
Original Article
View Cached Full Text

Cached at: 06/28/26, 03:59 AM

What’s the best model you can run on your @NVIDIAAI DGX Spark?

1× DGX Spark

  • ⁠Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s
  • DeepSeek v4 Flash REAP
  • ⁠Qwen 3.6 27b - 256k ctx 19, tok/s

2× DGX Sparks ← sweet spot!

  • DeepSeek v4 Flash - 1M ctx, 40-45 tok/s
  • Step-3.7-Flash - 256k ctx, image support, 30 tok/s

4× DGX Sparks

  • GLM 5.2 NVFP4 on all 4
  • 2× DeepSeek v4 Flash setups, each on 2× DGX Sparks.

Links and repos below

Similar Articles

DGX Spark agentic usage numbers

Reddit r/LocalLLaMA

A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.