Setting up of a 16xGB10 (DGX Spark) cluster
Summary
A user is setting up a 16-node DGX Spark cluster to run frontier open models locally, discussing networking tradeoffs between 200 and 100 Gbit/s.
Similar Articles
“The All Spark” Cluster: Upgrading from 16 - 36 DGX Sparks
The article details upgrading a personal homelab from 16 to 36 NVIDIA DGX Spark devices, forming a cluster for simultaneous AI inference tasks like model serving and media generation, chosen for cost-effectiveness and data sovereignty.
DGX Spark agentic usage numbers
A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.
@onusoz: 16x parallel Gemma-4-26B-A4B-NVFP4 runs 18 output tokens/s, aggregate 300 tok/s 1 DGX Spark with 128 GB unified memo…
@onusoz demonstrates running 16 parallel instances of NVIDIA's quantized Gemma-4-26B-A4B-NVFP4 model on a single DGX Spark with 128GB unified memory, achieving 300 tok/s aggregate, showcasing high concurrency without flashinfer.
@Tech2Wild: Running GLM-5.2 at home the FULL 744B, all 256 experts, UNPRUNED across 4× NVIDIA DGX Spark (GB10). 200K context · MTP …
A detailed recipe for running the unpruned GLM-5.2 model (744B parameters, 256 experts) across 4 NVIDIA DGX Spark nodes with 200K context, achieving up to 60.5 tok/s aggregate. Includes performance benchmarks, credits, and patches.
@exolabs: https://x.com/exolabs/status/2103617535765573959
This article is a handbook for the NVIDIA DGX Spark, a device designed for local AI inference, detailing its specifications, how to link multiple units for enhanced performance, and its optimization for running mixture-of-experts models.